Live data from Hacker News

Google search only has 60% of my content from 2006

tablix.org

61–70 of 168 posts

Re: Google search only has 60% of my content from 2006

#61
post #7
post #4

You'll see this effect from every search engine. They have no choice, there are a lot of sites with an infinite number of pages; so instead the number of pages they store per site depends on how important your site is, and they try to store your top N pages by relative importance.

I'm not sure I buy that they have no choice. For websites that literally have an infinite number of (dynamically generated) pages, sure, they could detect that and exclude them. But we're talking about unique, static pages here. And they don't even have to store the whole page, just the indexed info. I read this as, they could, but it's cheaper not to, and most people won't notice anyway.

I'm not sure that's true. How can one automatically determine whether a page is unique or static? As a trivial example, a URL path that accepts arbitrary strings and hashes them generates unique, immutable pages, but obviously cannot be crawled in its totality.

Re: Google search only has 60% of my content from 2006

#62
post #38

Earlier quoted context omitted.

Their A/B test told them to do it, without wondering if they should do it Basically their engagement numbers were better for a larger amount of people by making search engines counterintuitive for early adopters. We personally need a good robotic search engine that indexes like a robot. Everyone else needs a semi-sentient thing that makes many assumptions about what they want to see.

> Basically their engagement numbers were better for a larger amount of people by making search engines counterintuitive for early adopters. Which also makes sense ... if you present the "right" result immediately, the user visits one site and has completed whatever he sought to do. if you make him click through 10 pages, he has way more chances to see an interesting ad.

Good points although in Google’s case the first several results are ads and their main users cant differentiate and dont care even if they could, followed by amp pages by the most engaged webmasters optimizing for relevancy

That user wants fingerprint based ads and recent articles

Google is optimized for that

We are the only ones that want a “search engine”, a service distinctly good at indexing the known universe, instead of merely presenting the paid and compliant universe

Re: Google search only has 60% of my content from 2006

#63

Earlier quoted context omitted.

> only 1 in 3000 pages gets indexed ... we should see this ratio continue to degrade until this fundamental architecture is replaced. Content on the internet is growing exponentially. Processing power is not. Losing access to information is just one of the many sad implications of the death of Moore's law.

Simple Bloom filter, generated and published by site itself, can lower the curve a lot.

A bloom filter based on keywords, published on the site?

Re: Google search only has 60% of my content from 2006

#64
post #43

With all the talk about Google results not being satisfying anymore to a growing number of users, I'm surprised we haven't seen more sites pop up that would allow users to display the results of multiple search engines of their choosing either by mixing (eg all 1st results then all 2nd, etc) or by seeing them side by side... while stripping ads and cards and the like.

We had that back in the late 90s. I remember dogpile and Copernic.

Actually, I just checked them out, and it seems both of those are still alive.

Re: Google search only has 60% of my content from 2006

#65
post #49
post #13

Why does Google deeply index those useless telephone directory sites? Try searching for the impossible U.S. phone number "307-139-2345" and you'll see a bunch of "who called me?" or "reverse phone number lookup" sites. Virtually all of those sites are complete garbage. They make no attempt to collect numbers from telephone directories or from the web. They won't identify a number as being the main phone number for Di…

Actually it does a good job. When there is some data (like crowd sourced number reporting) on internet about the specific phone number you searched for, Google will show it in the top results You get a full page of non-sense results and ads/spams when the phone number you searched for is not known from any website (I guess)

That is my experience as well.

Re: Google search only has 60% of my content from 2006

#66
https://slashdot.org/comments.pl?sid=7132077&cid=49308245

From my short dystopian story, The Time Rift of 2100: How We lost the Future

"IN A SAD IRONY as to the supposed superiority of digital over analog --- that this whole profession of digitally-stored 'source' documentation began to fade and was finally lost. It had became dusty, and the unlooked-for documents of previous eras were first flagged and moved to lukewarm storage. It was a circular process, where the world's centralized search indices would be culled to remove pointers to things that were seldom accessed. Then a separate clean-up where the fact that something was not in the index alone determined that it was purgeable. The process was completely automated of course, so no human was on hand to mourn the passing of material that had been the proud product of entire careers. It simply faded."

"THEN SOMETHING TOOK THE INTERNET BY STORM, it was some silly but popular Game with a perversely intricate (and ultimately useless) information store. Within the space of six months index culling and auto-purge had assigned more than a third of all storage to the Game. Only as the Game itself faded did people begin to notice that things they had seen and used, even recently, were simply no longer there. Or anywhere. It was as if the collective mind had suffered a stroke. Were the machines at fault, or were we? Does it even matter? Life went on. We no longer knew much about these things from which our world was constructed, but they continued to work."

Re: Google search only has 60% of my content from 2006

#67
post #13

Why does Google deeply index those useless telephone directory sites? Try searching for the impossible U.S. phone number "307-139-2345" and you'll see a bunch of "who called me?" or "reverse phone number lookup" sites. Virtually all of those sites are complete garbage. They make no attempt to collect numbers from telephone directories or from the web. They won't identify a number as being the main phone number for Di…

https://www.google.com/search?hl=en&q=307%2D139%2D2345

You comment now comes up first, but the rest of the results all try to contact googlesyndication.com, so ads? Google will not exclude sites that literally give them money.

Re: Google search only has 60% of my content from 2006

#68
You used to be able to google a simple question, something that could be answered on the search page without having to click through. But since no one clicked on them, they stopped appearing after a few years. The only results were ones where the data was hidden and you had to click through.

Re: Google search only has 60% of my content from 2006

#69
post #7

Earlier quoted context omitted.

I'm not sure I buy that they have no choice. For websites that literally have an infinite number of (dynamically generated) pages, sure, they could detect that and exclude them. But we're talking about unique, static pages here. And they don't even have to store the whole page, just the indexed info. I read this as, they could, but it's cheaper not to, and most people won't notice anyway.

I'm not sure that's true. How can one automatically determine whether a page is unique or static? As a trivial example, a URL path that accepts arbitrary strings and hashes them generates unique, immutable pages, but obviously cannot be crawled in its totality.

> How can one automatically determine whether a page is unique or static?

They crawled it for years and it never changed? It is a blog post.

Post reply on HN