Live data from Hacker News

It seems that Google is forgetting the old web

stop.zona-m.net

21–30 of 311 posts

Re: It seems that Google is forgetting the old web

#21

This author is in his own little bubble and doesn't understand the vast amount of blog-repost spam that google has to deal with. The way their algorithm most likely deals with this is a mixture of domain rank + tenure... how long has this copy of this article existed on this domain, and can we be sure this is the original copy? The author says the article was removed in 2006 (" [...] posts, were not accessible anymor…

Yeah, but how DDG was able to show "the new original"?

Re: It seems that Google is forgetting the old web

#22
post #12

The assertion that this is because " indexing the whole Web is crushingly expensive, and getting more so every day " is a bit flawed. Since old content is very unlikely to be updated, it doesn't have to be re-crawled a lot. I'm certain Google has a score that tells it how often the content of a given site is likely to change. This argument of expense becomes even less durable when you consider that DuckDuckGo, a comp…

Crawling isn't the real problem, nor is the bulk storage for the crawled pages.

What do you do with these pages after you've crawled them? You need to build an index out of them, and serve that index out of some kind of low latency storage (DRAM, Flash). That makes increasing the index size very expensive. The index size has to be limited, and selecting the right pages to include in the index is thus a core quality feature for a search engine.

Re: It seems that Google is forgetting the old web

#25

This author is in his own little bubble and doesn't understand the vast amount of blog-repost spam that google has to deal with. The way their algorithm most likely deals with this is a mixture of domain rank + tenure... how long has this copy of this article existed on this domain, and can we be sure this is the original copy? The author says the article was removed in 2006 (" [...] posts, were not accessible anymor…

Part of the problem is that their algorithm has become weighted against blogs and personal websites.

> Rumors spread that large link pages (for surfing) might be considered “link farms” (and yes on SEO sites they were but these things eventually trickle down to little personal site webmasters too) so these started to be phased out. Then the worry was Blogrolls might be considered link farms so they slowly started to be phased out. Then the biggie: when Google deliberately filtered out all the free hosted sites from the SERP’s (they were not removed completely just sent back to page 10 or so of the Google SERP’s) and traffic to Tripod and Geocities plummeted. Why? Because they were taking up space in the first 20 organic returns knocking out corporate and commercial sites and the sites likely to become paying customers were complaining.

https://ramblinggit.com/2018/08/when-the-social-silos-fall/

SEO seems to have become a huge obstacle course that smaller websites can't play.

Re: It seems that Google is forgetting the old web

#26
post #22
post #12

The assertion that this is because " indexing the whole Web is crushingly expensive, and getting more so every day " is a bit flawed. Since old content is very unlikely to be updated, it doesn't have to be re-crawled a lot. I'm certain Google has a score that tells it how often the content of a given site is likely to change. This argument of expense becomes even less durable when you consider that DuckDuckGo, a comp…

Crawling isn't the real problem, nor is the bulk storage for the crawled pages. What do you do with these pages after you've crawled them? You need to build an index out of them, and serve that index out of some kind of low latency storage (DRAM, Flash). That makes increasing the index size very expensive. The index size has to be limited, and selecting the right pages to include in the index is thus a core quality f…

I'm having trouble imagining that Google would be more limited by the ratio of hardware power vs data size today than it was in the early days. If keeping the whole index in DRAM is now a requirement, then yes, I'd expect a hugely reduced overall dataset - but wouldn't that affect way more sites/pages than the comparatively few dropped historical records?

I still suspect that this whole thing is more about bias (and personalization, be it correct or incorrect) in the results.

Re: It seems that Google is forgetting the old web

#27

This author is in his own little bubble and doesn't understand the vast amount of blog-repost spam that google has to deal with. The way their algorithm most likely deals with this is a mixture of domain rank + tenure... how long has this copy of this article existed on this domain, and can we be sure this is the original copy? The author says the article was removed in 2006 (" [...] posts, were not accessible anymor…

Yesterday I noticed that Google Scholar forgot one of my articles from 2018, on arXiv. See: https://scholar.google.com/scholar?q=arXiv%3A1811.04960 Google Scholar is not the same as Google Search, which can still find it https://www.google.com/search?q=arXiv%3A1811.04960 For how long, I have no idea. The article was at the same link all the time and arXiv is very reputable.

Re: It seems that Google is forgetting the old web

#29
post #18

The post unflatteringly compares Google with DuckDuckGo. But doesn't DuckDuckGo use Google's technology?

No, it uses Bing AFAIK

Not only Bing, from their docs:

"In fact, DuckDuckGo gets its results from over four hundred sources. These include hundreds of vertical sources delivering niche Instant Answers, DuckDuckBot (our crawler) and crowd-sourced sites (like Wikipedia, stored in our answer indexes). We also of course have more traditional links in the search results, which we also source from a variety of partners, including Oath (formerly Yahoo) and Bing."

(Regarding search results + instant answers: https://help.duckduckgo.com/duckduckgo-help-pages/results/so...)

Post reply on HN