Live data from Hacker News

It seems that Google is forgetting the old web

stop.zona-m.net

71–80 of 311 posts

Re: It seems that Google is forgetting the old web

#71

This author is in his own little bubble and doesn't understand the vast amount of blog-repost spam that google has to deal with. The way their algorithm most likely deals with this is a mixture of domain rank + tenure... how long has this copy of this article existed on this domain, and can we be sure this is the original copy? The author says the article was removed in 2006 (" [...] posts, were not accessible anymor…

Here is an example: http://www.gnoosic.com/discussion/metallica__5.html No matter how you search for the content on Google, nothing comes up: https://www.google.com/search?q="Metallica+only+played+2+son... DuckDuckGo has it: https://duckduckgo.com/?q="Metallica+only+played+2+songs+fro... I checked the wayback machine and the content has constantly been on that url for over 10 years. This is the first example of an ol…

Interestingly, even though DuckDuckGo finds the post, Bing doesn't seem to.

Re: It seems that Google is forgetting the old web

#72
post #45

Personally I have similar experience, but other way around. Every time I try look for anything in Google almost always most of results are from 3-6 years ago unless I specifically specify I want results from last month / year / etc. And I not just talking about technical questions, but all kind of stuff include music, travel information and such. I not even sure when the last time google provide me the link to some n…

I don't think your "3-6 years ago" and "forgetting the old web" are incompatible. I've noticed the same - Google seems to gravitate to results from 2014ish, even when newer information is available, or when I'm searching about an event from far before that time.

Re: It seems that Google is forgetting the old web

#73

Earlier quoted context omitted.

Part of the problem is that their algorithm has become weighted against blogs and personal websites. > Rumors spread that large link pages (for surfing) might be considered “link farms” (and yes on SEO sites they were but these things eventually trickle down to little personal site webmasters too) so these started to be phased out. Then the worry was Blogrolls might be considered link farms so they slowly started to…

You're jumping from describing observable results to a state of mind or motive which you can't observe. > Then the worry was Blogrolls might be considered link farms so they slowly started to be phased out. Then the biggie: when Google deliberately filtered out all the free hosted sites from the SERP’s... That's all observable fact. Why? Because they were taking up space in the first 20 organic returns knocking out c…

I absolutely agree with you that whether Google is intentionally diabolical or not is up in the air. My reason for quoting Brad there is to succinctly recount a history where Google has been a menace (deliberate or not) to individual blogs and websites. Blog rolls were absolutely a great way to discover new blogs and were hardly “link farms” but were an incredibly valuable resource. (An equivalent to modern friend lists.)

Where I don’t agree with you is in the portrayal of the Web as largely comprised of link farms and “few legitimate pages”. I spend a lot of my time cataloging the hidden corners of the Web and it is mostly individuals working on their personal Web projects. Spam is simple to identify (much more so than ‘clickbait’) and many of the reasons people don’t read personal websites any more isn’t because interesting and mind-blowing projects on the Web are too rare. (I don’t have statistics to back this up, but I feel like they are more common on the Web than on social media.)

Re: It seems that Google is forgetting the old web

#74

Earlier quoted context omitted.

Here is an example: http://www.gnoosic.com/discussion/metallica__5.html No matter how you search for the content on Google, nothing comes up: https://www.google.com/search?q="Metallica+only+played+2+son... DuckDuckGo has it: https://duckduckgo.com/?q="Metallica+only+played+2+songs+fro... I checked the wayback machine and the content has constantly been on that url for over 10 years. This is the first example of an ol…

Interestingly, even though DuckDuckGo finds the post, Bing doesn't seem to.

It's on Yandex.

https://www.yandex.ru/yandsearch?text=Metallica%20only%20pla...

Re: It seems that Google is forgetting the old web

#75
post #56

Earlier quoted context omitted.

Not only Bing, from their docs: "In fact, DuckDuckGo gets its results from over four hundred sources. These include hundreds of vertical sources delivering niche Instant Answers, DuckDuckBot (our crawler) and crowd-sourced sites (like Wikipedia, stored in our answer indexes). We also of course have more traditional links in the search results, which we also source from a variety of partners, including Oath (formerly…

How current is that page? Yahoo used Google since 2015.

That just raises the question: To what extend does DuckDuckGo still depend on Bing?

I've asked before, and no one seems to be able to provide me with a source that states that DuckDuckGo is just Bing. Yet it comes up every time DuckDuckGo is mentioned in a positive light.

Re: It seems that Google is forgetting the old web

#76
post #48
post #27

Earlier quoted context omitted.

Yesterday I noticed that Google Scholar forgot one of my articles from 2018, on arXiv. See: https://scholar.google.com/scholar?q=arXiv%3A1811.04960 Google Scholar is not the same as Google Search, which can still find it https://www.google.com/search?q=arXiv%3A1811.04960 For how long, I have no idea. The article was at the same link all the time and arXiv is very reputable.

I also noticed that all our scholarly articles are gone from Google Scholar. The only thing there is our two highly cited books. https://scholar.google.ca/scholar?hl=en&as_sdt=0%2C5&q=site%... We've come to rely on Google too much, so much that if you are not on Google you don't exist. That's a problem with researchers that are looking for articles to cite.

Somebody starts a site for collecting "scholar dropouts"? An article qualifies as a scholar dropout if:

- it was previously available on Google Scholar

- it cannot be retrieved, or the search on Google Scholar gives a misleading result (for example it gives another article, as explained in [1])

Please help to make a list of scholar dropouts! Thank you.

[1] https://news.ycombinator.com/item?id=19604722 HN comment with evidence

[2] https://news.ycombinator.com/item?id=19604955 HN reply with another evidence

Re: It seems that Google is forgetting the old web

#77
post #37

I think this is not "forgetting the older web" but the algorithm simply penalizing things that look like they pretend to be from the past.

That would be interesting, because it seems like a lot of blogs pretend to be from the present by updating the post/last modified date frequently but leaving the content the same as several years ago.

Re: It seems that Google is forgetting the old web

#78
I have been noticing (what appear to be) truncated search results in Google for some time now. At first I thought that was because I was accessing Google through Startpage. But that's not the case.

Anyways, I find myself using Bing more and more often these days, because the search results dig more deeply into the 'obscure'.

I'm not at all upset by this. It seems to me that as Google's results are not completely satisfactory, more people will make use of various alternatives. Maybe one day, search will become decentralized again, somewhat like it was in the 1990s, when you regularly made use of many search engines, like Altavista, Lycos, Excite, and Yahoo.

I would imagine that there must still be metasearch sites out there somewhere that submit your query to several search engines. I need to find one again and would appreciate recommendations.

Re: It seems that Google is forgetting the old web

#80
post #65

The reason is simple: apparently they are not making money on old content searches.

And the reason they are not making money (I'm guessing) is because older web sites are much less likely to use one of Google's spying plugs, such as Analytics or Fonts. Or the ads themselves after all.
Post reply on HN