You'll see this effect from every search engine. They have no choice, there are a lot of sites with an infinite number of pages; so instead the number of pages they store per site depends on how important your site is, and they try to store your top N pages by relative importance.
I'm not sure I buy that they have no choice. For websites that literally have an infinite number of (dynamically generated) pages, sure, they could detect that and exclude them. But we're talking about unique, static pages here. And they don't even have to store the whole page, just the indexed info. I read this as, they could, but it's cheaper not to, and most people won't notice anyway.
Google search only has 60% of my content from 2006
61–70 of 168 posts
Re: Google search only has 60% of my content from 2006
#62Earlier quoted context omitted.
Their A/B test told them to do it, without wondering if they should do it Basically their engagement numbers were better for a larger amount of people by making search engines counterintuitive for early adopters. We personally need a good robotic search engine that indexes like a robot. Everyone else needs a semi-sentient thing that makes many assumptions about what they want to see.
> Basically their engagement numbers were better for a larger amount of people by making search engines counterintuitive for early adopters. Which also makes sense ... if you present the "right" result immediately, the user visits one site and has completed whatever he sought to do. if you make him click through 10 pages, he has way more chances to see an interesting ad.
That user wants fingerprint based ads and recent articles
Google is optimized for that
We are the only ones that want a “search engine”, a service distinctly good at indexing the known universe, instead of merely presenting the paid and compliant universe
Re: Google search only has 60% of my content from 2006
#63Earlier quoted context omitted.
> only 1 in 3000 pages gets indexed ... we should see this ratio continue to degrade until this fundamental architecture is replaced. Content on the internet is growing exponentially. Processing power is not. Losing access to information is just one of the many sad implications of the death of Moore's law.
Simple Bloom filter, generated and published by site itself, can lower the curve a lot.
Re: Google search only has 60% of my content from 2006
#64With all the talk about Google results not being satisfying anymore to a growing number of users, I'm surprised we haven't seen more sites pop up that would allow users to display the results of multiple search engines of their choosing either by mixing (eg all 1st results then all 2nd, etc) or by seeing them side by side... while stripping ads and cards and the like.
Actually, I just checked them out, and it seems both of those are still alive.
Re: Google search only has 60% of my content from 2006
#65Why does Google deeply index those useless telephone directory sites? Try searching for the impossible U.S. phone number "307-139-2345" and you'll see a bunch of "who called me?" or "reverse phone number lookup" sites. Virtually all of those sites are complete garbage. They make no attempt to collect numbers from telephone directories or from the web. They won't identify a number as being the main phone number for Di…
Actually it does a good job. When there is some data (like crowd sourced number reporting) on internet about the specific phone number you searched for, Google will show it in the top results You get a full page of non-sense results and ads/spams when the phone number you searched for is not known from any website (I guess)
Re: Google search only has 60% of my content from 2006
#66From my short dystopian story, The Time Rift of 2100: How We lost the Future
"IN A SAD IRONY as to the supposed superiority of digital over analog --- that this whole profession of digitally-stored 'source' documentation began to fade and was finally lost. It had became dusty, and the unlooked-for documents of previous eras were first flagged and moved to lukewarm storage. It was a circular process, where the world's centralized search indices would be culled to remove pointers to things that were seldom accessed. Then a separate clean-up where the fact that something was not in the index alone determined that it was purgeable. The process was completely automated of course, so no human was on hand to mourn the passing of material that had been the proud product of entire careers. It simply faded."
"THEN SOMETHING TOOK THE INTERNET BY STORM, it was some silly but popular Game with a perversely intricate (and ultimately useless) information store. Within the space of six months index culling and auto-purge had assigned more than a third of all storage to the Game. Only as the Game itself faded did people begin to notice that things they had seen and used, even recently, were simply no longer there. Or anywhere. It was as if the collective mind had suffered a stroke. Were the machines at fault, or were we? Does it even matter? Life went on. We no longer knew much about these things from which our world was constructed, but they continued to work."
Re: Google search only has 60% of my content from 2006
#67Why does Google deeply index those useless telephone directory sites? Try searching for the impossible U.S. phone number "307-139-2345" and you'll see a bunch of "who called me?" or "reverse phone number lookup" sites. Virtually all of those sites are complete garbage. They make no attempt to collect numbers from telephone directories or from the web. They won't identify a number as being the main phone number for Di…
You comment now comes up first, but the rest of the results all try to contact googlesyndication.com, so ads? Google will not exclude sites that literally give them money.
Re: Google search only has 60% of my content from 2006
#68Re: Google search only has 60% of my content from 2006
#69Earlier quoted context omitted.
I'm not sure I buy that they have no choice. For websites that literally have an infinite number of (dynamically generated) pages, sure, they could detect that and exclude them. But we're talking about unique, static pages here. And they don't even have to store the whole page, just the indexed info. I read this as, they could, but it's cheaper not to, and most people won't notice anyway.
I'm not sure that's true. How can one automatically determine whether a page is unique or static? As a trivial example, a URL path that accepts arbitrary strings and hashes them generates unique, immutable pages, but obviously cannot be crawled in its totality.
They crawled it for years and it never changed? It is a blog post.
Re: Google search only has 60% of my content from 2006
#70Arguably, search is such a vital function of modern society that it could be considered a public good and seized on the principal of eminent domain.