Live data from Hacker News

Google Memory Loss

tbray.org

511–520 of 552 posts

Re: Google Memory Loss

#511
post #508
post #206

Earlier quoted context omitted.

You know, that actually makes a lot of sense. Recently, I attempted to bid on some long-tail keywords in AdWords for some targeted ads, but unfortunately they don’t have the “search volume” to qualify.

You mean you cannot bid on keywords with very low search volume ?

Yes Adwords will prevent you from targeting low volume/long tail combinations and recommend higher volume more vague and simple keyword/combinations.

Re: Google Memory Loss

#512

Earlier quoted context omitted.

I think the biggest irony is that the web allows for more adoption of long-tail movements than ever before, and Google has gotten significantly worse at turning these up. I assume this has something to do with the fact that information from the long tail is substantially less searched for than stuff within the normal bounds. Google wants you to be mainstream now. If everyone thinks the same and wants the same things…

I see quite the opposite inceintive for Google. If you are a very eccentric individual and they know those quirks, they have a huge competitive advantage in targeting ads to you vs. some bulk radio broadcast ad etc.

I see quite the opposite incentive for Google. If you are a very eccentric individual and they know those quirks, they have a huge competitive advantage in targeting ads to you vs. some bulk radio broadcast ad etc.

However, the observation of this article, and my observation as well, is that Google isn't currently capable of parsing very individual quirks. Rather, Google is able to place you into one of a number of highly conformist boxes. They don't have to understand you as an individual. They just have to 'box' you more effectively than their competitors.

There is nothing in the market or otherwise emergent in the nature of data and such categorization which fundamentally motivates Google to be able to parse anyone's quirks or understand the essence of a scene or artistic movement. If Google can gain a competitive advantage by creating a number of honeytrap doppelgangers which draw people away from the long tail and sequester them into un-creative, imitative, and highly conformist boxes, then so much the better for them.

https://www.youtube.com/watch?v=R9OHn5ZF4Uo

http://lesswrong.com/lw/l8/conjuring_an_evolution_to_serve_y...

In much the same way, I find that recommendation engines come up with annoying pale imitations of bands/musicians I like. I also wonder why authoritarianism seems to spread so effectively across social media, and why certain authoritarian movements seem to get such ready support from within Google and various social media companies. It's because, as a product, conformist/authoritarian screechers are more easily herded, replicated, categorized, and packaged than real individuals who think for themselves and apply principles.

Re: Google Memory Loss

#513

Earlier quoted context omitted.

I can understand "firce" and "fprce" or even "f9rce" but "furce" is, on a standard QWERTY keyboard, two keys away so more unlikely to be a typo.

Do Google do spelling correction based on letter locality on the expected user keyboard? Never seen any corrections that would suggest that, often wondered why not.

There used to be a service that would, given a query, search eBay for listings that matched it or common misspellings based on nearby keys. I wonder if any of that logic has made its way into modern autocorrect explicitly or if it would be gathered implicitly through studying what users actually correct.

Re: Google Memory Loss

#514
post #273
post #211

Earlier quoted context omitted.

No, they don't but I know what you are referring to. Most of the time I get the result: No results found for [phrase] Results for [phrase] (without quotes) ..but the thing is that I know websites containing [phrase] exist. Often many of them and they aren't "dark web" either. Google used to be able to find them but no longer is. This gets more confusing because sometimes it does work. Namely if you look for phrases w…

I had also wondered whether its ability to match exact phrases has degraded. It would make sense, since keeping individual words on a distributed index is a lot easier than keeping a long phrase. But I had no way of telling, not having benchmarked it in the past.

Can they not index bigrams or trigrams, then chain together index hits? E.g. "It never rains in december" would hit on "It never rains", "never rains in", and "rains in december". Any result that hits on all of the indexes is not guaranteed to hit on the entire phrase but it would be a good candidate for the top result. The longer the phrase, the more likely a candidate hitting all necessary index phrases would match the exact phrase. This would at least put a limit on how large the indexes need to get.

Further, if they retain copies of the full text in their database they could do a filtered scan of the documents that hit on all subphrases to guarantee exact match. I could see that having too much of a performance impact at scale though.

In any case the dumbing down of Google search over the last few years is immensely frustrating to me.

Re: Google Memory Loss

#515
post #327

Earlier quoted context omitted.

I start feeling like the web is being de-optimized for nerds & super-users

Nothing wrong with that. You generally shouldn't optimize applications for power users. Software should be usable.

Good software is when both kinds of users are pleased. As Alan Kay said "Simple things should be simple, complex things should be possible.". But it's probably the hardest thing in UX design to make something that is simple and powerful at the same time - most things sadly end up being only one of the two.

Re: Google Memory Loss

#516
post #189

Hmm. I just tried to reproduce this with old posts of my own and couldn't. I picked random phrases from five early 2006 blog posts that get basically no traffic and searched for them: "I had been playing the accordion Davy lent to Rosie during winter break" "The language they're using is not that different from the one I wrote PlayGUI to use" "I've been playing a decent amount of music lately, mostly guitar and piano…

Actually, rssing.com has nothing (directly) to do with your RSS feed. It seems to be a content scraper and mirroring/archiving "service," if I'm feeling charitable. And it looks like a site that wants to redirect users to its copy of users' content in order to get ad revenue if I'm feeling less-than-charitable.

In fairness to Google, that's always been a problem, and even in the good-old-days in which keyword-based searches were more effective, there were content aggregators that would copy the entire contents of phpBB-style bulletin boards (and USENET newsgroups) in order to rehost them and get clicks.

On the one hand, I want to say that it's precisely the sort of SEO/spammy practice that Google should be deprioritizing in search results. On the other hand, sometimes these copies/mirrors of content are the only extant copies of content when an original blog goes away. Although the motivations of the owners of these sorts of sites may not be as pure as that of archive.org, the result for the searcher is equivalent: the desired information is found even if it's only a rehosted copy.

Re: Google Memory Loss

#517
post #330

Earlier quoted context omitted.

Startup idea: a service that will let you search your inbox. Aka google for searching. Seriously, this is egregious. You rely on your email provider to accurately search your inbox - some emails are important business, tax, and legal documents that are relevant for years, even decades. Or at least be fucking transparent about the fact that you are not really searching all emails. I know Gmail is a free service and in…

Thunderbird (and I'm guessing most of the offline, true email clients) has this built-in.

It's funny, in discussions about webmail vs local email client, people often say "why would I want a local client anymore, webmail is all I need".

Well... this is why you might want it. Your data under your control. Your choice of tools.

If you're using GMail and Google decides to turn GMail to crap, well, bad luck.

Re: Google Memory Loss

#518
post #505

Earlier quoted context omitted.

I agree that having that option would make sense. But arguably the default should be correction, because most users are probably making typos.

Fine but I am logged in, if it were an option I could fix it once and carry on.

Yes, I totally agree.

Re: Google Memory Loss

#519
post #296

Earlier quoted context omitted.

Which means the corpus is broken. But regular people rarely care about correct spelling, in my experience, and so I doubt corpus maintainers will care either...

You're imagining that people who make NLP corpora actually vet the text going into them? I dream of a world where people can be convinced to care that much. I'm not even talking about the scenario you suggest of filtering for proper word usage, I'm talking about filtering at all . The corpora used for popular word embeddings are full of weird nonsense text (in the case of word2vec) or autogenerated awfulness like spa…

I mean, nobody expects the engineers to manually read through everything, but if the quality of the input text is significant for the quality of the autocorrect (or whatever other application you're using machine learning with), you kind of have to make sure the input is pretty good... You could for example choose datasets which is expected to contain mostly correct grammar and spelling (such as Wikipedia, books, etc.) rather than datasets which is expected to contain mostly incorrect grammar and spelling.

Or don't use a machine learning model. I honestly don't care, just don't automatically turn a correct "its" into an incorrect "it's".

Re: Google Memory Loss

#520
post #213

Earlier quoted context omitted.

That was just a random example of a phrase. I made it up on the fly.

Not sure what you expect then.

He isn't complaining that the specific phrase "It never rains on a Wednesday in Rockshire" to turn up any results. He used it as an example of the kind of phrase which, even though they exist somewhere on the web, won't be found with Google.
Post reply on HN