Live data from Hacker News

Google Memory Loss

tbray.org

321–330 of 552 posts

Re: Google Memory Loss

#321

I wouldn't be surprised if this issue is as mysterious to Google staff as it is to you. Google Search no longer runs a clearly defined algorithm to find search results. It is a collection of AI systems that are trained continuously on a variety of data. There is probably no human alive who fully understands how Google makes decisions about which results to return and how to rank them. They just understand how to prov…

That's unfortunate, because their simpler system worked better.

I don't think so. A few years ago, Google worked much better when you knew how to phrase queries. I often helped family members to find something online, just by rephrasing their original query.

Today, it doesn't matter how you build your query, Google returns good results in any case. That also means that you can't search for specific info by phrasing queries differently but for the vast majority of people it makes life much easier.

Re: Google Memory Loss

#322

Earlier quoted context omitted.

Try switching between languages. I work in English, my wife is Italian, my colleagues are french speaking, I am Spanish and have Catalan friends, and I live in Germany. I gave up on predictive keyboards long ago.

SwiftKey seems to manage three languages (Finnish, Swedish, English) simultaneously pretty well. No need to explicitly switch language either, it figures out the current language as you type.

Ditto on the SwiftKey recommendation: I have German, English and French always at my disposal.

Re: Google Memory Loss

#323

Earlier quoted context omitted.

No, I don't. Can you help clarify it for me? Search engines crawling millions of sites each with---on average---a few MB of data distributes cost globally. Extracting terabytes of index data from a single search engine's repository consolidates the cost on the back of that repository's bandwidth provision. These are not symmetrical cost structures.

Our git repository went down when crawlers decided to index it

But probably not Google. The google crawler is very careful and stops as soon as they encounter higher error rates. Bing appears to do the same.

Re: Google Memory Loss

#324

I'm wondering if rackless Ruth, Google's bean-counter-in-chief, is behind all these. At Bing, bean counters often had calculated that if you cut down index to half (after certain size), you reduce half of the cost but don't lose half of the revenues. So there is a sweet spot where you can maximize revenue if you are willing to let go few demanding customers. When quarterly results needs a little push, everything is a…

Did you really intend to call her "rackless Ruth" or was the slur an accidental misspelling of reckless? Because pointless and crude name-calling like that really detracts from any valid points you might have brought up.

Re: Google Memory Loss

#325

I've noticed this many times too, particularly recently, and I call it "Google Alzheimer's" --- what was once a very powerful search engine that could give you thousands (yes, I've tried exhausting its result pages many times, and used to have much success finding the perfect site many dozens of pages deep in the results) of pages containing nothing but the exact words and phrase you search for has seemingly degraded…

The worse is when it gives you results containing synonyms of your query words (and even highlights them in the little description under the link). Like, dude, I used this word for a reason

Re: Google Memory Loss

#326
post #212

Theory: Because google has local data-centers all over the world (that's why it's much faster than the competition, and google suggest works so well), indexes must be maintained at all of them. Because google has so many, this is a significant expense, and to keep costs down, they reduce the indexes to what is profitable.

Good theory, maybe that's why they push the "personalization" and localization of search engine so aggressively. I personally don't feel that Google search quality has degraded very much if at all. It's true that I rarely get new sites outsite the echo chamber or few 1000 popular sites, but to me 99% of the time they give me relevant and useful information I am looking for. To be honest Google search results on avera…

I think they push localisation because it makes sense for the user. In Europe, queries for at least 10 countries will usually be answered by the same data center and are still localised. The main reason why I still often use Google instead of DDG is to find localised content where you can't localise by language (e.g. finding UK specific issues).

Re: Google Memory Loss

#327

Earlier quoted context omitted.

It’s a bit funny and sad how Google has both managed to seem overly clever to the point of being clumsily useless, and at the same time, continues to offer the original search experience that people pine for but no one knows or can be bothered with: add “intext:” before keywords and it’s the experience you’re wishing for.

I just tried a few of my old "dry" queries (as in, things I've always wanted to look for but can't seem to get good results) and didn't notice much difference... Furthermore, after a total of 3 "Next Page"'s, I got IP-banned with a CAPTCHA. I've had similar experiences recently with "site:" and the other colon-operators, so not entirely unexpected, but still immensely infuriating. It's almost like any real attempt to…

I start feeling like the web is being de-optimized for nerds & super-users

Re: Google Memory Loss

#328

Earlier quoted context omitted.

This won’t answer all of questions but the measures you’re looking for are called ‘recall’ and ‘precision‘: - recall: number of relevant documents retrieved / number of relevant documents - precision: number of relevant documents in result set / number of documents in result set

Yeah you know, it's funny, the last time I worked on question-answering code, we were trying really hard to find algorithms that could improve a particular metric (F-score, a synthetic agglomeration of precision and recall) ... I don't remember hearing very many conversations at all about whether we were measuring the right thing. Given a query like [site:tbray.org "rock n roll animal"], and knowing that the 1 releva…

"Modern Information Retrieval" by Baeza-Yates / Ribeireo Neto a few years ago used to be a good standard work.

I'm not sure though how well it's kept up in terms of aspects like real-time search and graph search, both of which are fairly recent developments.

Re: Google Memory Loss

#329

Earlier quoted context omitted.

> has seemingly degraded into an approximation of a search engine that has knowledge of only very superficial information, will try to rewrite your queries and omit words (including the very word that makes all the difference...) I think the biggest irony is that the web allows for more adoption of long-tail movements than ever before, and Google has gotten significantly worse at turning these up. I assume this has s…

This same AI effect can be seen in the Android keyboard, where _properly_ spelled words will be replaced after typing another word or two because it's been determined to be more likely what you want. It's infuriating.

And they can't even fix a simple typo like this after years of usage - "th8s 8s example tex6". Come on Google, you made that UI, you know that i and 8 are next to each other, you have a database of correct words and probably of typical errors and typos, wtf?! (you even know how to correct "wft" to "wtf" and can correct simple word with number typo).

Re: Google Memory Loss

#330
post #17

I've convinced myself that this happens in gmail / hangouts history search too. It'll very confidently tell you that here are the only six results for your search term going back to the beginning of time, but if you go and manually dig up something that you know is there from ten years ago, then all of a sudden there are seven results the next time you search for the same term. I haven't done this methodically, and I…

Startup idea: a service that will let you search your inbox. Aka google for searching. Seriously, this is egregious. You rely on your email provider to accurately search your inbox - some emails are important business, tax, and legal documents that are relevant for years, even decades. Or at least be fucking transparent about the fact that you are not really searching all emails. I know Gmail is a free service and in…

Thunderbird (and I'm guessing most of the offline, true email clients) has this built-in.
Post reply on HN