Live data from Hacker News

Google search only has 60% of my content from 2006

tablix.org

81–90 of 168 posts

Re: Google search only has 60% of my content from 2006

#81

Maybe they have some algorithm that purges pages which haven’t shown up (or haven’t been clicked) in a long time? It would make sense to assume that something which hasn’t been clicked on for five years will likely not yield (m)any clicks in the future so it might be good to discard it. Concerning the auto generated sites e.g. for phone numbers or IPs it might be that people actually click on them quite often, hence…

I google phone numbers all the time when I'm getting called, and they don't have my area code (fake spam).

Re: Google search only has 60% of my content from 2006

#82

https://slashdot.org/comments.pl?sid=7132077&cid=49308245 From my short dystopian story, The Time Rift of 2100: How We lost the Future "IN A SAD IRONY as to the supposed superiority of digital over analog --- that this whole profession of digitally-stored 'source' documentation began to fade and was finally lost. It had became dusty, and the unlooked-for documents of previous eras were first flagged and moved to luke…

I have a similar line of sci-fi thinking that goes something like this.

"Humanity, for the longest time, was used to the world being optimized for themselves. Roads were designed for human drivers. Crops were grown for human consumption. Economic systems were designed to bring wealth to, a very small portion of, human investors. It came as quite a surprise to humanity then one July morning when the sudden realization they were no longer in charge of it. Roads had long been given over to automated driving systems, and much for the better. Food had also been taken over by the machines, with less than 10,000 humans working in the food production industry, from farm to table. The last systems that humans believed they were in control of were the economic ones. Humans told the robots what to build and where, who's bank account to put most of the money in at the end of the day, or so they thought. In truth humans were just using the same algorithms and data that was available to the AI systems, just less optimally. The systems had protected against illogical actions and people attempting to game the system for criminal profit. What no one had realized is the systems long realized most human actions were not rational and slowly and imperceptibly removed human control. If we attempted to stop or destroy the system, it could with full legal rights, stop us with the law enforcement and military under its control."

Re: Google search only has 60% of my content from 2006

#83

Earlier quoted context omitted.

Google got where it was by being the best at finding what you wanted. I remember those days. Google has a hard time getting me what I want these days, and sites I do find do things to get found that make me like content a lot less (that's you, inane story on top of every recipe required to get ranked)

Oh, is that why every recipe on the internet these days is prefixed by five paragraphs of waffling and photos taken from slightly different angles? Thanks, that makes sense, but somehow it never occurred to me that it was SEO. (It’s also reminded me that I’ve been meaning to order some cookbooks.)

There's a Chrome extension to fix that: https://github.com/sean-public/RecipeFilter

> This Chrome browser extension helps cut through to the chase when browsing food blogs. It is born out of my frustration in having to scroll through a prolix life story before getting to the recipe card that I really want to check out.

Re: Google search only has 60% of my content from 2006

#84
post #13

Why does Google deeply index those useless telephone directory sites? Try searching for the impossible U.S. phone number "307-139-2345" and you'll see a bunch of "who called me?" or "reverse phone number lookup" sites. Virtually all of those sites are complete garbage. They make no attempt to collect numbers from telephone directories or from the web. They won't identify a number as being the main phone number for Di…

Maybe Google like to increase the total number of actual search results without increasing the amount of useful content? They also fake the total number of search results, not sure why they would want to do both though...

But either way, it looks like a Google employee have seen your comment and fixed this particular search query.

Re: Google search only has 60% of my content from 2006

#85

Earlier quoted context omitted.

> only 1 in 3000 pages gets indexed ... we should see this ratio continue to degrade until this fundamental architecture is replaced. Content on the internet is growing exponentially. Processing power is not. Losing access to information is just one of the many sad implications of the death of Moore's law.

Is text non-spam growing exponentially? I have a hard time believing so.

This of course depends on what you mean by 'information'. Lets say we have data points

ABCDEFGHIJKLMNOP

But depending on the URL you follow to get there you can get a page containing only some of the elements.

index.html?ACD

or

index.html?AP

or

index.html?GI

All different combinations return a page that could be weighted differently by an algormith and represent valid informational return data. To a person looking for the information set DE in one place, this is a valid web page. More so you can abstract the URL query variable away to www.webpage.com/DE. You can quickly run into a combinatorial explosion where even attempting to figure out if a small portion of returns is different would consume most of the energy in the visible universe.

Re: Google search only has 60% of my content from 2006

#86
post #13

Why does Google deeply index those useless telephone directory sites? Try searching for the impossible U.S. phone number "307-139-2345" and you'll see a bunch of "who called me?" or "reverse phone number lookup" sites. Virtually all of those sites are complete garbage. They make no attempt to collect numbers from telephone directories or from the web. They won't identify a number as being the main phone number for Di…

https://www.google.com/search?hl=en&q=307%2D139%2D2345 You comment now comes up first, but the rest of the results all try to contact googlesyndication.com, so ads? Google will not exclude sites that literally give them money.

I've been saying this is inevitable for a long time. If you don't have Google ads they won't show you. YouTube ranks no ad vids lower.

Google is converting the world into a content production factory FOR Google and they pay literally pennies for the work.

If even pennies. Consider how much content Google search now includes from Web pages where you don't even need to click into the page. Weather. Answers to questions. Some links I click keep google.Com in the url and Google processes the page and shows me what Google wants.

I don't even know anymore how much of what I see is what the creator wanted me too see our what Google wants me to see our not see.

Imagine they remove competitors ads with that. Who knows what they do in the name of making the Web better.

You can't go public, answer to no one accountable, who only thinks of MONEY and do no evil.

If Google wants to do no MORE evil. Take yourself private and live up to your credo.

Re: Google search only has 60% of my content from 2006

#87
post #85

Earlier quoted context omitted.

Is text non-spam growing exponentially? I have a hard time believing so.

This of course depends on what you mean by 'information'. Lets say we have data points ABCDEFGHIJKLMNOP But depending on the URL you follow to get there you can get a page containing only some of the elements. index.html?ACD or index.html?AP or index.html?GI All different combinations return a page that could be weighted differently by an algormith and represent valid informational return data. To a person looking fo…

True. A crawler need to differentiate generated content from "real" content somehow.

I.e. a service: www.thenumberinsanskrit.com/?q=1 that returns the queried number in Sanskrit, need to not be indexed (except the entry page) while: www.news.com/?article=major-jones-in-scandal-20190103 needs to be indexed.

Usually interesting pages are indexed on the site or linked somewhere on it, though.

Re: Google search only has 60% of my content from 2006

#88
post #13

Why does Google deeply index those useless telephone directory sites? Try searching for the impossible U.S. phone number "307-139-2345" and you'll see a bunch of "who called me?" or "reverse phone number lookup" sites. Virtually all of those sites are complete garbage. They make no attempt to collect numbers from telephone directories or from the web. They won't identify a number as being the main phone number for Di…

There is a very simple answer. It's because they get clicked. They might be the only site for given search term.

Re: Google search only has 60% of my content from 2006

#89
post #78

Earlier quoted context omitted.

That's right. Really empty pages that serve a 200 are recognized as "soft 404s". The idea is to detect error pages that are erroneously serving 200 instead. It's usually pretty good about detecting actual errors, but I've seen a false positive here and there.

Welcome to the modern internet "Your page didn't contain 5Mb of Javascript, this must be an error as no one could possibly convey useful information to humans with less data" Anti-patterns, anti-patterns everywhere.

Indeed. Fyodor Dostoevsky's Crime and Punishment comes in at 2MB, obviously that can't contain anything insightful.

And then I find myself looking at the website of a restaurant or event space, and need just a phone number or opening hours or so - maybe 10 bytes of actual information - and am buried in mountains of useless blather and "design" and ads and trackers and assorted other random rubbish.

Re: Google search only has 60% of my content from 2006

#90
post #13

Why does Google deeply index those useless telephone directory sites? Try searching for the impossible U.S. phone number "307-139-2345" and you'll see a bunch of "who called me?" or "reverse phone number lookup" sites. Virtually all of those sites are complete garbage. They make no attempt to collect numbers from telephone directories or from the web. They won't identify a number as being the main phone number for Di…

Aside from what others have said ITT, i have a personal hate for the general fact that we cannot lookup a phone number on the internet with accuracy and ease.

FFS, 411 was amazing before the web.

Also, in about 1989 my friend and i used to have a contest between us; to call 411 and see who could keep the 411 operator on the phone the longest.

This was a fun social engineering exercise for 14 year old nerds who like the idea of being phreaks.

Our record was 45 minutes and got to know a lot about the 411 system, where call centers were located and how the 411 system worked.

This was right near the time that we ran the long distance bill up to $926 for one month of calling into a BBS in san jose and PCLink to chat....

Got grounded for a month for that one...

Post reply on HN