Earlier quoted context omitted.
Requiring users to know what sites they want in advance somewhat defeats the purpose of a search engine, no?
since sites are so desperate to be indexed, doesn't it seem better to put the onus on them to announce themselves? it would be great if dns registries publshed public keys .. maybe they do in newer schemes?
Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
431–440 of 492 posts
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#432Earlier quoted context omitted.
I feel your pain. Two workarounds when Google gets it wrong are to put the term in quotation marks, or to enable Verbatim mode in the toolbelt. (I know various people have come up with ways to add "Google Verbatim" as a search engine option in their browser, or use a browser extension to make Verbatim enabled by default.) Disclaimer: I work on Google search.
Try this, go to Google and type in "eggzackly this". Two results not containing "eggz" at all. Two results containing "eggzackly this" Two results containing "eggzackly" but missing "this". Google Search is broken. It no longer does what it's directed, it just takes a guess. I suspect part of this is because someone decided that "no results found" was the worst possible result a search engine could give.
Therefor I don't see how your last sentence is the explanation (there are results), I've also happened to land on no results found sometimes with overly precise quoted queries (for coding errors mostly IIRC). But it is annoying that it doesn't seem stricktly enforced even when you want it to.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#433Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
> Hardware costs are too high I want to say - you don’t know what are talking about. But, it’ll be rude. Hardware is much cheaper and powerful now compared to 2005.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#434Earlier quoted context omitted.
What would trolling the entire web look like?
It would look like a modern search engine with innovative technology offerings like Advanced Mobile Pages.
AMP is the perfect way to troll websites into making shitty versions of their content, for no real reason other than just because you feel like it. And then when you’re satisfied with your trolling you just abandon the standard.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#435Earlier quoted context omitted.
Isn't it's familiarity to early Google a side-effect of the early Internet being text-heavy sites in the first place rather than a similarity in the search engine? Unless I am misunderstanding your site's intent, even if you reach the dream engine you are trying to achieve, I won't be using it to search answers for coding questions on SO, how-tos for car repair, sites to stream movies, governmental page for X need, t…
I guess it depends on what you are looking for on the Internet I guess. Right now the biggest problem with Marginalia is that it has a fairly uneven quality level. For some queries it's absolutely incredible. For others, it doesn't really provide much useful results at all. I do think it's possible to even that out a considerable bit, to make it more viable for general queries. It's never going to be able to answer e…
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#436Earlier quoted context omitted.
I get what you mean, but part of the whole initial appeal of Google was that it gave much more relevant results initially than Altavista or the other options. That was why Google put in the audacious "I'm feeling lucky" button.
>"I'm feeling lucky" button My brain got so used to ignoring it I completely forgot it's a thing. I'm also unclear what it does? On an empty request, it gets me to their doodles page and with text in the box, gets me to my account history landing page.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#437Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#438Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
> You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our pre…
For me the experienced quality of Google search results gave have dropped massively since 2008, despite (and maybe even because of) all their new parameters.
When someone says this someone else usually immediately says it is because of web spam and black hat SEO.
But black hat SEO doesn't explain why verbatim doesn't work for many of us.
Black hat SEO doesn't explain why double quotes doesn't work.
Black hat SEO doesn't explain why there is no personal blacklists so all those who hate pintrest can blacklist them.
Black hat SEO probably also doesn't explain why I cannot find a unique strings in open source repos and instead get pages of not exactly webspam but answers to questions I didn't ask.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#439Earlier quoted context omitted.
I think people also have an inflated recollection of how good Google actually was back in 2005. Back then Google was only going up against indexes and link-rings, not 2021 Google/Bing/DDG/etc.
I've been using Altavista at that time, every now and then switching to Northern Light. Everything else was abysmal. Google blew them out of the water in terms of speed, quality, simplicity, unclutterdness and everything else. I can't remember ever retraining muscle memory so fast when switching to Google. So, no, Google has been great then and apart from people actively working against the algorithm is still good no…
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#440Earlier quoted context omitted.
> You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our pre…
I think people also have an inflated recollection of how good Google actually was back in 2005. Back then Google was only going up against indexes and link-rings, not 2021 Google/Bing/DDG/etc.
I've been pointing this out for at least close to a decade.
I know since I bothered to screenshot and blog about it in 2012.
I'll admit mistakes happened back then too, but they were more forgivable like keyword stuffing on unrelated pages. Back then Google were on our side and removed those as fast as possible.
Today however the problem isn't that someone hss stuffed the keyword into an unrelated page but that Google themselves mix a whole lot of completely irrelevant pages into the results, probably because some metrics go up when they do that.
Thinking about it it seems logical that for a search engine that practically speaking has monopoly both on users and as mattgb points out - tonsome degree also on indexing - serving the correct answer first is just dumb: if they can keep me going between their search results and tech blogs with their ads embedded one, two or five times extra that means one, two or five times more ad impressions.
Note that I'm not necessarily suggesting an grand evil master plan here, only that end-to-end metrics will improve as long as there is no realistic competition.