Live data from Hacker News

Google Memory Loss

tbray.org

21–30 of 552 posts

Re: Google Memory Loss

#21
post #17

I've convinced myself that this happens in gmail / hangouts history search too. It'll very confidently tell you that here are the only six results for your search term going back to the beginning of time, but if you go and manually dig up something that you know is there from ten years ago, then all of a sudden there are seven results the next time you search for the same term. I haven't done this methodically, and I…

Sounds like they don't index the full corpus

Re: Google Memory Loss

#22
post #17

I've convinced myself that this happens in gmail / hangouts history search too. It'll very confidently tell you that here are the only six results for your search term going back to the beginning of time, but if you go and manually dig up something that you know is there from ten years ago, then all of a sudden there are seven results the next time you search for the same term. I haven't done this methodically, and I…

This has happened to me with labels before, too. I'll do a search for all things that have a label and are in my inbox, and then archive them. I'll then go back to my inbox, and see that it missed something with that label. If I then repeat the search, I get zero results, even though it has the label, is in my inbox, and I can go back and find it. It's extremely frustrating.

Re: Google Memory Loss

#23
I would bet this is one of those more subtle long-term effects that nobody really saw coming... when Google refocused search with an eye towards commercial results, I imagine it deprioritized a lot of the older, more innocent informational content lying around

Re: Google Memory Loss

#24
Nobody can index the whole web. Even a single site in the form of

    Homepage of Joe Infinity
    You are on page 
    ">Next Page
can not be completely indexed. A search engine will crawl it to some depth based on many factors. Age might be one of them. There is no way to index 'everything' on the web.

Re: Google Memory Loss

#25
post #24

Nobody can index the whole web. Even a single site in the form of Homepage of Joe Infinity You are on page ">Next Page can not be completely indexed. A search engine will crawl it to some depth based on many factors. Age might be one of them. There is no way to index 'everything' on the web.

That's really not relevant to this article. The author is not talking about crawling and indexing the entire web (although he mentions the "whole web" once, that's clearly not what he means). He is wondering why old pages -- pages that used to be in Google's index -- are no longer showing up in SERPs even when using appropriately-targeted long-tail queries.

Re: Google Memory Loss

#26
post #7

I would not be surprised if google still has the data. Not sure how google handles things internally. However, google needs to pull up the results fast. So they might have 4 billion results with the word "water" in it. They only make tiny portion of that available. So if I type the words "Hot water" google it looks at the subset of pages with words "Hot" and the word "Water" So google must pull the pages that have bo…

Intersections are another thing that Google search doesn't do properly anymore. If I search for something like lkasdfjer samsung galaxy s8 it just gives me matches for samsung galaxy s8 and ignores the first word. When I do searches like this, I do it for a reason and don't want matches that lack some of the search terms.

I've found if I put the keyword in double-quotes then it makes the keyword required in the search

Re: Google Memory Loss

#27

This is definitely a thing. Many sites when they update take down old content that was getting few views. Many times that content is irretrievably lost. This shows the value of actually grabbing content that you plan to use or hope to refer to in the future, rather than merely bookmarking it. And it also underscores the value of the Internet Archive.

> This shows the value of actually grabbing content that you plan to use or hope to refer to in the future, rather than merely bookmarking it. And it also underscores the value of the Internet Archive.

Yes. Everyone should install the Wayback Machine plugin and click the "save page now" whenever they find something useful or interesting:

Chrome: https://chrome.google.com/webstore/detail/wayback-machine/fp...

Firefox: https://addons.mozilla.org/en-US/firefox/addon/wayback-machi...

I hate hitting an unarchived dead-ends when doing research, so I'm trying to do my part to prevent it. Many page I've archived had never been archived before I saved them.

Re: Google Memory Loss

#28
post #5

https://www.google.com/search?q="we+were+watching+the+Democr... I can find the article this person is referring just by looking at a string of it. So the article _is_ indexed. Not sure what he is referring to.

I would be surprised if it has not been reindexed thanks to the link from this article.

Re: Google Memory Loss

#29
post #25
post #24

Nobody can index the whole web. Even a single site in the form of Homepage of Joe Infinity You are on page ">Next Page can not be completely indexed. A search engine will crawl it to some depth based on many factors. Age might be one of them. There is no way to index 'everything' on the web.

That's really not relevant to this article. The author is not talking about crawling and indexing the entire web (although he mentions the "whole web" once, that's clearly not what he means). He is wondering why old pages -- pages that used to be in Google's index -- are no longer showing up in SERPs even when using appropriately-targeted long-tail queries.

For the same reason 'Joe Infinity page 1234567' would not be found anymore. Google thinks its not relevant enough to keep it indexed. Yes, it is debatable what is relevant enough and what isn't. But everyone who indexes 'the web' has to decide what to keep and what not. Nobody can store 'everything'.

Also it's not as easy as just keeping everything that ever was in the index in there. Then searchengines would link to noexisting urls most of the time. Most URLs have a short lifespan. Links rot pretty fast.

Re: Google Memory Loss

#30
This explains a number of times I've been unable to find old articles/forums/what have you, even when fairly certain I recall most or all of their titles. This may finally be enough for me to move to DuckDuckGo, as the quantity of information published longer ago increases, and the information I may wish to reference becomes increasingly difficult to locate.
Post reply on HN