Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

141–150 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#141

Brave's new search engine seems to work pretty well. Have been using it as my primary for about 10 days, and so far, I've only had to revert to Google once, and when I did the results were chock full of spam.

The nice thing about Brave Search is that they're trying to create an index completely independent from Bing/Google, and they seem to be trying to innovate on ways to get there as well with their Web Discovery Project[0], unlike DuckDuckGo. They've announced Brave Search will get ads soon, with a premium version without ads, which I think is acceptable given the costs of running an independent index sustainably.

[0]: https://brave.com/privacy/browser/#web-discovery-project

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#142
I feel like we are at the low point or even losing the battle between search engines and SEO spam. Maybe it is time for the Yahoo-style curated directory to return? We seem to be getting a microcosm of this with the awsome-* GitHub lists and Gemini with its near-nonexistent search.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#143

Some people try: https://www.mojeek.com/ https://fireball.com/ https://search.brave.com/

Mojeek founder story here: https://blog.mojeek.com/2021/03/to-track-or-not-to-track.htm... No-tracking and independent from the start. Now at 4.6 billion pages with own infrastructure and IP. Went to market in 2020 with contextual ads and API. Self-disclosure: CEO

HN is wild: 30m after something is mentiond, the CEO chimes in.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#144
post #124

Earlier quoted context omitted.

It does operate on a scale and principle fairly similar to early 2000s google, so the comparison isn't that far off, but yeah, it's quite some way before it's viable for general search. Dunno if I'll ever get there, but it does consistently seem to get better so who knows.

Isn't it's familiarity to early Google a side-effect of the early Internet being text-heavy sites in the first place rather than a similarity in the search engine? Unless I am misunderstanding your site's intent, even if you reach the dream engine you are trying to achieve, I won't be using it to search answers for coding questions on SO, how-tos for car repair, sites to stream movies, governmental page for X need, t…

I guess it depends on what you are looking for on the Internet I guess.

Right now the biggest problem with Marginalia is that it has a fairly uneven quality level. For some queries it's absolutely incredible. For others, it doesn't really provide much useful results at all. I do think it's possible to even that out a considerable bit, to make it more viable for general queries. It's never going to be able to answer every query, but it probably could answer a lot more than it does.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#145

Why don't you want personalized results? If I search for "subaru service" I want to find Austin Subaru, not Thorp Subaru in Cape Town.

Why didn't you just search "austin subaru service"? If you want a query narrowed down by location, that's your job to say so. Sure, it feels great when the engine guesses something like that correctly -- but it comes out worse overall for the plentiful cases where you have to try to compensate for it guessing wrong.

Why should I have to do all that work? I want the machine to do it for me.

I can only think of examples where I want personalization. What's an example query where it interferes?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#146
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

Regarding the gatekeeper problem: it's a wild guess but maybe if there was a way to involve users by organizing distributed scraping just for the sake of building a decent index, I'm sure many of them would help.

yes, large proxy networks are potential solutions. but they cost money, and you are playing a cat and mouse game with turing tests, and some sites require a login. furthermore, people have tried to use these to spider linkedin (sometimes creating fake accounts to login) only to be sued by microsoft who swings the CFAA at them. so you start off with an intellectual desire to make a nice search engine and end up getting sidetracked into this pit of muck and having microsoft try to put you in jail. and, no, i'm not the one microsoft was suing.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#147

Earlier quoted context omitted.

Of course, but stemming is a fairly basic technique in NLP, as is POS-tagging. NLP is not machine learning.

Stemming is not meaningfully a natural language processing technique, any more than arithmetic is a technique of linear equations.

Is it not the processing of natural language?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#148
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

Not sure if you're looking for feedback, but the News search could use some work, I searched for "Ethiopia" and almost all of the articles were unrelated to Ethiopia except for the existence of some link somewhere on the page.

Your general web search seems pretty good, although I've just given it a casual glance. I think your News search could be improved by just filtering the general search results for News-related content, since the "Ethiopia" content I get there is certainly Ethiopia-related.

In any case, an interesting product, I'll try to keep an eye on it.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#149
post #77
post #38

Earlier quoted context omitted.

Trusted curators is a dangerous dependency

Trusted consumers are better. The original page-rank algo was organic and bottom-up. But now it's the person not the page. Businesses compete for interaction not inbound links. So if you can make a modern page-rank that follows interaction instead of links and isn't a walled garden then I'd invest.

I could make that work, but what do you mean by "walled garden" in this context?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#150
post #81
post #22

No mention of DDG in the comments? Is there a reason I'm not seeing or it's just not the preferred alt-search on HN? Seems to have been working fine for me when I struggle to get past the funnels and content mills on Google.

DDG doesn't have their own index (they're getting their results from Bing) so not really relevant to this question.

I.. didn't know that. However, trying it just now in incognito I don't get the same results[0] (some different links, and most re-ordered). Is Duck repurposing Bing's results? I've tested with "how to get rich", a great bait for bad content (try it on Google without an adblocker, if you dare).

[0]: https://pastebin.com/xC45hL1i

Post reply on HN