Live data from Hacker News

It seems that Google is forgetting the old web

stop.zona-m.net

31–40 of 311 posts

Re: It seems that Google is forgetting the old web

#32

This author is in his own little bubble and doesn't understand the vast amount of blog-repost spam that google has to deal with. The way their algorithm most likely deals with this is a mixture of domain rank + tenure... how long has this copy of this article existed on this domain, and can we be sure this is the original copy? The author says the article was removed in 2006 (" [...] posts, were not accessible anymor…

They even have a tool just for this: canonical URLs. It lets websites specify which version is the source/canonical version and avoids the old copy indexing.

Re: It seems that Google is forgetting the old web

#33

While it's become impossible to browse the wider Web with Google, it's getting a bit easier elsewhere. A few helpful search engines: * https://millionshort.com/ * https://wiby.me/ * https://pinboard.in/search/ A recent movement to build personal Yahoo!-style directories: * https://href.cool/ (my own project) * https://indieseek.xyz/ * https://districts.neocities.org/ * https://the.dailywebthing.com/ The above resourc…

How do I get updates on the groups project?

Wiby seems amazing - the first three surprise me links were a human powered ornithopter, lego maniacs and a guide to knife throwing. Thank you for sharing!

Re: It seems that Google is forgetting the old web

#34
post #22
post #12

The assertion that this is because " indexing the whole Web is crushingly expensive, and getting more so every day " is a bit flawed. Since old content is very unlikely to be updated, it doesn't have to be re-crawled a lot. I'm certain Google has a score that tells it how often the content of a given site is likely to change. This argument of expense becomes even less durable when you consider that DuckDuckGo, a comp…

Crawling isn't the real problem, nor is the bulk storage for the crawled pages. What do you do with these pages after you've crawled them? You need to build an index out of them, and serve that index out of some kind of low latency storage (DRAM, Flash). That makes increasing the index size very expensive. The index size has to be limited, and selecting the right pages to include in the index is thus a core quality f…

The index only has to be on "low latency storage" if low-latency results to any query are required. While that's definitely true of the modal "Google Search", most of these queries for "long-tail, old content" as discussed in the OP don't really need that sort of quick response.

Re: It seems that Google is forgetting the old web

#36

This author is in his own little bubble and doesn't understand the vast amount of blog-repost spam that google has to deal with. The way their algorithm most likely deals with this is a mixture of domain rank + tenure... how long has this copy of this article existed on this domain, and can we be sure this is the original copy? The author says the article was removed in 2006 (" [...] posts, were not accessible anymor…

Yeah, but how DDG was able to show "the new original"?

DDG isn't only using Google, it uses other search engines too.

Re: It seems that Google is forgetting the old web

#38

I have noticed that searching for exact quotes seems to have been broken on Google for a few years. But only minimally broken. And I've had no idea how to reason with it. This article completely corresponds with problems I've encountered with searching for results on StackOverflow or software documentation sites; it's especially perplexing that "site:..." combined with exact quotes does not work for many cases. Googl…

I too noticed that for some queries, Google is becoming really, really unwieldy.

I can't recall the exact search term, but I kept looking for some site I visited some time ago, and no combination of words could get it to actually find the actual site. I finally just gave up and found it in my browser history.

Re: It seems that Google is forgetting the old web

#39
I would not be surprised that if along the way, googles search results were optimized for income of ad revenue, over other metrics - knowingly or not.

I to have complained about the search results of google going down hill. I was told "I was just to technical".

Whatever the case, the web is not the same as it was in the early 2000s and it really is sucking if you are wanting to search for something.

Re: It seems that Google is forgetting the old web

#40

While it's become impossible to browse the wider Web with Google, it's getting a bit easier elsewhere. A few helpful search engines: * https://millionshort.com/ * https://wiby.me/ * https://pinboard.in/search/ A recent movement to build personal Yahoo!-style directories: * https://href.cool/ (my own project) * https://indieseek.xyz/ * https://districts.neocities.org/ * https://the.dailywebthing.com/ The above resourc…

We're among those that believe there is room in the search space. We're building out our hyper local product search service city by city at https://attic.city. It's meant to fill the void with regard to smaller, non-chain stores that Google shopping seems to focus on.
Post reply on HN