Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

431–440 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#431

Earlier quoted context omitted.

Requiring users to know what sites they want in advance somewhat defeats the purpose of a search engine, no?

since sites are so desperate to be indexed, doesn't it seem better to put the onus on them to announce themselves? it would be great if dns registries publshed public keys .. maybe they do in newer schemes?

Certificate Transparency (CT) Logs are this.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#432
post #332

Earlier quoted context omitted.

I feel your pain. Two workarounds when Google gets it wrong are to put the term in quotation marks, or to enable Verbatim mode in the toolbelt. (I know various people have come up with ways to add "Google Verbatim" as a search engine option in their browser, or use a browser extension to make Verbatim enabled by default.) Disclaimer: I work on Google search.

Try this, go to Google and type in "eggzackly this". Two results not containing "eggz" at all. Two results containing "eggzackly this" Two results containing "eggzackly" but missing "this". Google Search is broken. It no longer does what it's directed, it just takes a guess. I suspect part of this is because someone decided that "no results found" was the worst possible result a search engine could give.

Googling that with the brackets I get results containing "eggzackly this" ranked 3, 4, 6 (your comment) and 7 whereas the others contain just eggzackly (or with the 'this' preceded by punctuation as you mention).

Therefor I don't see how your last sentence is the explanation (there are results), I've also happened to land on no results found sometimes with overly precise quoted queries (for coding errors mostly IIRC). But it is annoying that it doesn't seem stricktly enforced even when you want it to.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#433
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

> Hardware costs are too high I want to say - you don’t know what are talking about. But, it’ll be rude. Hardware is much cheaper and powerful now compared to 2005.

the complexity of the search algorithm has also increased substantially since 2005 And, in 2005, a billion page index was pretty big. Now it's closer to 100 billion.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#434
post #308

Earlier quoted context omitted.

What would trolling the entire web look like?

It would look like a modern search engine with innovative technology offerings like Advanced Mobile Pages.

Wow, you’re right. Trolling the entire web would involve an organization that carries considerable authority whose decisions can impact every member of the web.

AMP is the perfect way to troll websites into making shitty versions of their content, for no real reason other than just because you feel like it. And then when you’re satisfied with your trolling you just abandon the standard.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#435
post #124

Earlier quoted context omitted.

Isn't it's familiarity to early Google a side-effect of the early Internet being text-heavy sites in the first place rather than a similarity in the search engine? Unless I am misunderstanding your site's intent, even if you reach the dream engine you are trying to achieve, I won't be using it to search answers for coding questions on SO, how-tos for car repair, sites to stream movies, governmental page for X need, t…

I guess it depends on what you are looking for on the Internet I guess. Right now the biggest problem with Marginalia is that it has a fairly uneven quality level. For some queries it's absolutely incredible. For others, it doesn't really provide much useful results at all. I do think it's possible to even that out a considerable bit, to make it more viable for general queries. It's never going to be able to answer e…

Basically I understand Marginalia's proposition as a search engine focused on retrieving text-heavy/long form content. Unless I misunderstand it's intent, that can't replace a generalist engine (nor does it have to) as not every search request will lends itself to long form texts. I guess that's the only point I was going for (I do feel the old-Google sentiment has got more to do with the state of the web than the engine, but am out of my league for a proper opinion), and it certainly wasn't a jab at it - I'm thankful for your neat website and will be looking forward to see it get even better over time! Maybe it is somewhat uneven, but it is nonetheless great at finding thoughtful pieces written on subjects XYZ and surfacing more obscure/personal websites.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#436
post #427

Earlier quoted context omitted.

I get what you mean, but part of the whole initial appeal of Google was that it gave much more relevant results initially than Altavista or the other options. That was why Google put in the audacious "I'm feeling lucky" button.

>"I'm feeling lucky" button My brain got so used to ignoring it I completely forgot it's a thing. I'm also unclear what it does? On an empty request, it gets me to their doodles page and with text in the box, gets me to my account history landing page.

It automatically redirects to the first search result.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#438
post #256
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

> You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our pre…

>but we probably don't notice how well they work when they give us the result we expect in the first three links.

For me the experienced quality of Google search results gave have dropped massively since 2008, despite (and maybe even because of) all their new parameters.

When someone says this someone else usually immediately says it is because of web spam and black hat SEO.

But black hat SEO doesn't explain why verbatim doesn't work for many of us.

Black hat SEO doesn't explain why double quotes doesn't work.

Black hat SEO doesn't explain why there is no personal blacklists so all those who hate pintrest can blacklist them.

Black hat SEO probably also doesn't explain why I cannot find a unique strings in open source repos and instead get pages of not exactly webspam but answers to questions I didn't ask.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#439
post #321

Earlier quoted context omitted.

I think people also have an inflated recollection of how good Google actually was back in 2005. Back then Google was only going up against indexes and link-rings, not 2021 Google/Bing/DDG/etc.

I've been using Altavista at that time, every now and then switching to Northern Light. Everything else was abysmal. Google blew them out of the water in terms of speed, quality, simplicity, unclutterdness and everything else. I can't remember ever retraining muscle memory so fast when switching to Google. So, no, Google has been great then and apart from people actively working against the algorithm is still good no…

I think the parent's point was that people say Google 2005 >> Google 2021, but it's pretty hard to make this comparison in an objective way. No doubt Google 2005 was way better than other offerings around at the time.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#440
post #256

Earlier quoted context omitted.

> You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our pre…

I think people also have an inflated recollection of how good Google actually was back in 2005. Back then Google was only going up against indexes and link-rings, not 2021 Google/Bing/DDG/etc.

> I think people also have an inflated recollection of how good Google actually was back in 2005.

I've been pointing this out for at least close to a decade.

I know since I bothered to screenshot and blog about it in 2012.

I'll admit mistakes happened back then too, but they were more forgivable like keyword stuffing on unrelated pages. Back then Google were on our side and removed those as fast as possible.

Today however the problem isn't that someone hss stuffed the keyword into an unrelated page but that Google themselves mix a whole lot of completely irrelevant pages into the results, probably because some metrics go up when they do that.

Thinking about it it seems logical that for a search engine that practically speaking has monopoly both on users and as mattgb points out - tonsome degree also on indexing - serving the correct answer first is just dumb: if they can keep me going between their search results and tech blogs with their ads embedded one, two or five times extra that means one, two or five times more ad impressions.

Note that I'm not necessarily suggesting an grand evil master plan here, only that end-to-end metrics will improve as long as there is no realistic competition.

Post reply on HN