Live data from Hacker News

It seems that Google is forgetting the old web

stop.zona-m.net

131–140 of 311 posts

Re: It seems that Google is forgetting the old web

#131
post #37

I think this is not "forgetting the older web" but the algorithm simply penalizing things that look like they pretend to be from the past.

Agreed. This article commits the error that I see all too often online: a variation of Halon's razor where people will attribute something they don't like to deliberate malice rather than seeking some other explanation.

I would be surprised if google is actually "forgetting" older web pages. I would think that the pages were still in the index but newer sites are just favored over older results.

Re: It seems that Google is forgetting the old web

#132

This author is in his own little bubble and doesn't understand the vast amount of blog-repost spam that google has to deal with. The way their algorithm most likely deals with this is a mixture of domain rank + tenure... how long has this copy of this article existed on this domain, and can we be sure this is the original copy? The author says the article was removed in 2006 (" [...] posts, were not accessible anymor…

Here is an example: http://www.gnoosic.com/discussion/metallica__5.html No matter how you search for the content on Google, nothing comes up: https://www.google.com/search?q="Metallica+only+played+2+son... DuckDuckGo has it: https://duckduckgo.com/?q="Metallica+only+played+2+songs+fro... I checked the wayback machine and the content has constantly been on that url for over 10 years. This is the first example of an ol…

From what I can tell, there's no links anywhere on the site to that particular page, you have to know the exact term and search it: http://www.gnoosic.com/discussion/

How is Google supposed to find that out?!

Re: It seems that Google is forgetting the old web

#133

This author is in his own little bubble and doesn't understand the vast amount of blog-repost spam that google has to deal with. The way their algorithm most likely deals with this is a mixture of domain rank + tenure... how long has this copy of this article existed on this domain, and can we be sure this is the original copy? The author says the article was removed in 2006 (" [...] posts, were not accessible anymor…

Part of the problem is that their algorithm has become weighted against blogs and personal websites. > Rumors spread that large link pages (for surfing) might be considered “link farms” (and yes on SEO sites they were but these things eventually trickle down to little personal site webmasters too) so these started to be phased out. Then the worry was Blogrolls might be considered link farms so they slowly started to…

It's not Google's fault this time.

The problem is that Blogspam is now a (legitimate) industry much bigger than Google can manage.

Google Search became a playground for marketing firms to dump content made by low-paid freelancers with algorithmically chosen keywords, links and headers. It's SEO on large scale. Everything is monitored via analytics and automatically posted to Wordpress. Every time Google tweaks its algorithm to catch it, they're able to A-B test and then change thousands of texts all at once.

Personal blogs can't even dream about competing with that.

In fact, those companies are actively competing with personal blogs by themselves: via tools like SEMRush and social media monitoring, they know which blogs are trending and use their tools to produce copycat content re-written by freelancers and powered by their SEO machine.

I know a startup that is churning 10 thousand blogposts per day on clients blogs, each costing from 2 to 5 dollars for a freelancer to write according to algorithmically defined parameters.

Just wait until they get posts written via OpenAI-style machine learning: the quality will be even lower.

Not only that: there's no need for black hat SEO anymore. Blogposts from random clients have links to others clients blogs, and it is algorithmically generated in order to maximize views and satisfy Google's algorithm. They have a gigantic pool of seemingly unconnected blogs to link to, so why not use it.

The irony is that companies buy this kind of blogspam to skip paying AdSense. Why pay when you can get organic search results? So not only they're damaging the usefulness of the SERP, they're directly eating Google's bottom line. These blogs also have ZERO paid advertising inside them, since they're advertising themselves.

That's the reason Bing, DuckDuckGo and Yandex still have "old web" results.

That puts Google in a very difficult position and IMO they're not wrong to fight it.

Re: It seems that Google is forgetting the old web

#134

Earlier quoted context omitted.

The web is fine, and search is fine. It's specifically Google search that's being destroyed by spam. It's odd to put forward the hypothesis that DuckDuckGo is now better at search (aggregation) than Google is at search. But that seems to be where we have landed.

I think it may be a simple consequence of the fact that Google Search is increasingly less of a searching engine and more of an answering engine. I think Google has been explicit about this (I may be wrong, but I seem to remember thinking about this because Google themselves said it). Essentially, I believe, they are no longer concerned about being a way to navigate all the material found on the internet. Instead, th…

In that case, they have a branding problem, and should rename Google Search to Google Answers.

Re: It seems that Google is forgetting the old web

#135

Earlier quoted context omitted.

Interestingly, even though DuckDuckGo finds the post, Bing doesn't seem to.

It's on Yandex. https://www.yandex.ru/yandsearch?text=Metallica%20only%20pla...

Not showing up here. Not even if I add quotes.

Re: It seems that Google is forgetting the old web

#136

While it's become impossible to browse the wider Web with Google, it's getting a bit easier elsewhere. A few helpful search engines: * https://millionshort.com/ * https://wiby.me/ * https://pinboard.in/search/ A recent movement to build personal Yahoo!-style directories: * https://href.cool/ (my own project) * https://indieseek.xyz/ * https://districts.neocities.org/ * https://the.dailywebthing.com/ The above resourc…

Thanks! for sharing links. wiby in particular is amazing!

Re: It seems that Google is forgetting the old web

#137
post #38

I have noticed that searching for exact quotes seems to have been broken on Google for a few years. But only minimally broken. And I've had no idea how to reason with it. This article completely corresponds with problems I've encountered with searching for results on StackOverflow or software documentation sites; it's especially perplexing that "site:..." combined with exact quotes does not work for many cases. Googl…

I too noticed that for some queries, Google is becoming really, really unwieldy. I can't recall the exact search term, but I kept looking for some site I visited some time ago, and no combination of words could get it to actually find the actual site. I finally just gave up and found it in my browser history.

Yeah, I made similar experiences. Also to dive into some random topic, Google is not really helpful and just suggests only really obvious things first like Wikipedia.

On the other hand, for day to day work at least for me it is still indispensable. Googling a random exception out of an unexpected stack trace works far better than with DDG for instance.

Re: It seems that Google is forgetting the old web

#138
It's not just the old web. Google's results for political content or content with political implications have become increasingly strange in the past couple years. I noticed this when trying to recall the story of a false rape accusation that occurred on my college campus when I was a student. This saga consists of three articles: a) initial reporting of the reported rape](https://news.cornell.edu/stories/2012/09/cornell-police-inve...) b) [a police statement that they had irrefutable video evidence that the initial report was false ](https://cornellsun.com/2012/11/28/cornell-police-report-of-a...) c) [an article defending the school's policies on sexual assault in light of a and b](https://cornellsun.com/2013/02/25/after-false-report-cornell...). All three of these articles are impossible to find unless you search for exactly the right thing.

For example, if you search for "Cornell Police: Report of Attempted Rape on Campus Was False" on Google (the exact title of the second article), you find it. But, if you search for any variation (e.g. "Cornell Police: False Report of Attempted Rape on Campus") - the later articles (b and c) are impossible to find.

I've never been able to find a satisfactory explanation for why the discoverability on these articles has been turned to zero. I think that there are some odious websites devoted to covering false rape accusations, and these articles may be inheriting the low reputation of the publications that link to them. Or, perhaps the hosting website (student newspaper) is doing something to de-rank the stories. In all cases though, it seems wrong that in response to a query "cornell trolley bridge false rape accusation 2012", Google's top results would be "Why false rape accusations are rarer than you think"

Re: It seems that Google is forgetting the old web

#139

When web directories like Yahoo lost out to web search engines like Google, we lost something crucial. While search is good for answering questions you know how to ask, browsing was exploratory and led us to know what we didn't know. When it comes to learning a complex topic like mathematics, this kind of serendipity was very useful. There are some amazing resources on the web, but googling won't let you discover tho…

'Awesome' directories have really given a nice resurgence to the Yahoo! style of organization. I think the problem with Yahoo! is that it simply got too big to be used as a directory. Niche directories are where it's at. (Also: Reddit wikis, which often are used similarly.)

> Niche directories are where it's at.

Agreed. And if we came up with a standard format for such lists you could make them searchable and we could end up with distributed searchable curated indexes that are not centrally controlled. And that is quite compelling compared to centralized fully algorithmic search systems run by mega-corps.

Re: It seems that Google is forgetting the old web

#140
post #106
post #42

Earlier quoted context omitted.

So as with all other open systems, spam is destroying the web.

It seems like the opposite actually -- spam is destroying Google. They're so big that it's worth blackhats spending significant resources to game their algorithm. That induced them to implement a spam filter which is now discarding the ham along with the spam. Which means that smaller search engines that aren't being targeted by spammers are now giving better results. That is a major long-term problem for Google if t…

IME, it's not blackhats anymore causing the problem. It's (legitimate, but shady) marketing agencies and startups handling thousands of customers and with deep pockets to do SEO research.
Post reply on HN