Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

261–270 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#261
post #83

Earlier quoted context omitted.

I get that for general-purpose searches this is a good idea, but it would be nice if there was an easy way to disable this when you know you don't want it - for example, for most programming searches, if I type SomeAPINameHere the most relevant results will always be those that include my search term verbatim. I don't need Google to helpfully suggest "Did you mean Some API Name Here?", which will virtually always ret…

I feel your pain. Two workarounds when Google gets it wrong are to put the term in quotation marks, or to enable Verbatim mode in the toolbelt. (I know various people have come up with ways to add "Google Verbatim" as a search engine option in their browser, or use a browser extension to make Verbatim enabled by default.) Disclaimer: I work on Google search.

Both of these options are disappointing, in my experience. Verbatim mode seems weirdly broken sometimes (maybe it's overly strict), and quoting things is rarely enough to convince Google that you really want to search for exactly that thing and not some totally different thing that it considers to be a synonym.

One porridge is too hot and the other is too cold. I know Google could find a happy compromise here if it wanted to. In fact, I bet there's some internal-only hacked-together version that works this way and actually gives an acceptable user experience for the kind of people who have shown up to this thread to show their dissatisfaction.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#262
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

> Cloudflare (owned in part by Google)

Please elaborate. Is there a special relationship between Cloudflare and Google?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#263
post #256
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

> You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our pre…

I think people also have an inflated recollection of how good Google actually was back in 2005.

Back then Google was only going up against indexes and link-rings, not 2021 Google/Bing/DDG/etc.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#264

Earlier quoted context omitted.

And if you really need to, DDG !bangs[0] make a search as simple as "!g mother google help me". The keyword thing is also available in Firefox as a browser feature, and elsewhere I'm sure, but nevertheless, it makes switching to DDG easier. (Plus I can directly go to the wiki page by using "!w", "!gm" for google maps, etc.) [0] https://duckduckgo.com/bang

The only bang I use is !gvb since DDG doesn't support verbatim searches.

Is this the same as enclosing the terms in quotes and using the !g bang?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#265
post #38

Earlier quoted context omitted.

Perhaps trolling the entire web is not useful today? I’d love a search engine where I can whitelist sites or take an existing whitelist from trusted curators.

Trusted curators is a dangerous dependency

That’s why you don’t make it a hard dependency and let people curate their own list of taste makers. They can share and exchange info about who good taste makers are and good one might even charge for access to exclusive flavors.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#266
post #256
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

> You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our pre…

Well if the result didn't appear in the first 5-10 pages, it's probably not in the index.

You can see it with other search engines. I challenge you to come up with a Google query for which a first-page result won't be seen within the first 10 pages of Bing results for the same query.

(Bonus points if that result is relevant).

There's only so much tweaking that personalization and other heuristic can do.

But if something is missing from them index, that's it.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#267
post #82

Earlier quoted context omitted.

Heh, I guess you mean "trawling" - trolling the entire web is something very different :)

"Trolling" is fine, see e.g. https://grammarist.com/usage/trawl-troll/#:~:text=Troll%20fo... .

Well, no, it's not fine.

See e.g. the source you linked, which explains the difference.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#268

Earlier quoted context omitted.

Plus, it scales less well than pure algorithmic search. This fight already happened, with a much smaller internet.

It works really, really well for libraries. Research libraries (and research librarians) are phenomenally valuable. I've missed them any time I'm not at a university. Both curators and algorithms are valuable. This goes for finding books, for finding facts and figures, for finding clothes, for finding dishwashers, and for pretty much everything else. I love the fact that I have search engines and online shopping, but…

+book stores. Curators can use algorithms to help them curate… Google’s SE is taking signals from poor curators imo.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#269
post #256
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

> You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our pre…

I assume the author has the ability to search the index to see if your preferred Google result is even indexed.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#270
post #16

Early 2000s google index ran in a garage. The current google index has dedicated power stations. It's a bit like the car industry - you could run a startup from your garage in the early days but you need titanic amounts of capital to compete now thanks to vertical integration. Major governments and billionaires can compete but everybody else is locked out of the market (most "startups" use bings index).

I was thinking about exactly that. If they used simpler index would they be getting better results? There's not a lot of selective pressure so they just keep adding to the index algorithm.
Post reply on HN