Live data from Hacker News

Google Memory Loss

tbray.org

531–540 of 552 posts

Re: Google Memory Loss

#531
post #522
post #233

So from a busi­ness point of view, it’s hard to make a case for Google in­dex­ing ev­ery­thing, no mat­ter how old and how ob­scure. I don't get that line of thought, somehow people have starting defend lack of quality as something expected or reasonable. The whole point of going to google is to find stuff, that includes "boring" old and obscure stuff that won't sell ads. But that is part of the deal, if google don't…

Google's job is to make money, and to grow and make ever increasing amounts of money, not to find stuff. If Google decides that it's worth it to lose some of its customers in order to reduce costs, then it's fair game for them to do so. Maybe a competitor will come in and steal marketshare from Google by filling the hole Google leaves behind when they increasingly make changes which annoy a subset of users (duckduckg…

What makes you believe this is a rational decision? Even if you reduce everything to "must make money"?

This hasn't anything to do with capitalism. It is pure greed. Companies willingly do anything for a slight increase in revenue even if they willingly acknowledge that it will cost them ten times as much in the (not so) long term.

There is nothing about capitalism that says you must be a colossal idiot, that's just a consequence of a poisonous culture where employees don't give a crap about the company but only focus on their own career.

Re: Google Memory Loss

#532

I'm wondering if rackless Ruth, Google's bean-counter-in-chief, is behind all these. At Bing, bean counters often had calculated that if you cut down index to half (after certain size), you reduce half of the cost but don't lose half of the revenues. So there is a sweet spot where you can maximize revenue if you are willing to let go few demanding customers. When quarterly results needs a little push, everything is a…

Interesting to note that even this post apparently is not indexed in Google: Do the search "site:https://news.ycombinator.com "rackless ruth"" and find 0.

Re: Google Memory Loss

#533
post #280

Earlier quoted context omitted.

4 years ago was a very different time. Software quality in general continues to drop. I’m not sure if that’s due to increasing incompetence in the software engineering workforce (unlikely, but possible), malice (more unlikely), or apathy (most likely). When wages don’t grow over 10 years, what incentive is there to write the best software you can?

I suspect the cause is a little more subtle (and terrifying). The majority of people don't give a shit. Correct spelling and obscure searches are not even on their radar, it's not a part of their reality. Don't let the comments here fool you -- it is a very specific, picky, technical crowd that frequents HN. The voice of "those who care" has always been a minority, although it used to matter more, simply because peop…

> The voice of "those who care" has always been a minority

To add to this, people who cares most powerfully had probably switched to alternative, open source, software solutions. This leaves the remaining group with less "care" on average so fewer would complain. Kinda like evaporative cooling.

Re: Google Memory Loss

#534

Earlier quoted context omitted.

Nothing wrong with that. You generally shouldn't optimize applications for power users. Software should be usable.

Good software is when both kinds of users are pleased. As Alan Kay said "Simple things should be simple, complex things should be possible.". But it's probably the hardest thing in UX design to make something that is simple and powerful at the same time - most things sadly end up being only one of the two.

I agree, but I didn't say that good software doesn't please its users. I said you generally shouldn't optimize for power users and that software should be usable.

Re: Google Memory Loss

#536

Earlier quoted context omitted.

If this comment is what teaches me I've been expected to do that, I'm going to throw my toys out of the pram. I've specifically googled for instructions on how I might be expected to use the swipe-style keyboards, and turned up nothing.

As far as I know it's in the tutorial they insist you do upon first enabling swiping edit: Don't actually see a tutorial in the app. Maybe I'm confused with another app such as Swype, but the same technique seems to apply to all.

The default keyboard on my Android didn't have any such tutorial. I did look in the app and online.

Re: Google Memory Loss

#537
post #519
post #296

Earlier quoted context omitted.

You're imagining that people who make NLP corpora actually vet the text going into them? I dream of a world where people can be convinced to care that much. I'm not even talking about the scenario you suggest of filtering for proper word usage, I'm talking about filtering at all . The corpora used for popular word embeddings are full of weird nonsense text (in the case of word2vec) or autogenerated awfulness like spa…

I mean, nobody expects the engineers to manually read through everything, but if the quality of the input text is significant for the quality of the autocorrect (or whatever other application you're using machine learning with), you kind of have to make sure the input is pretty good... You could for example choose datasets which is expected to contain mostly correct grammar and spelling (such as Wikipedia, books, etc…

Wouldn't late edition books with only corrected text be better, proofread, edited, proofread, edited, ... Google have millions of them they've assumed copyright of. Surely there's enough text there. Do they really just use random website text?? Nearly every news story I read has errors and they have style guides, trained writers, editors, etc..

Do publishers sell their published text as a mass for use in AI/ML? Like 1000 books, no images or frontispiece, etc., possibly jumbled by sentence/para/page.

Re: Google Memory Loss

#538

I'm wondering if rackless Ruth, Google's bean-counter-in-chief, is behind all these. At Bing, bean counters often had calculated that if you cut down index to half (after certain size), you reduce half of the cost but don't lose half of the revenues. So there is a sweet spot where you can maximize revenue if you are willing to let go few demanding customers. When quarterly results needs a little push, everything is a…

Interesting to note that even this post apparently is not indexed in Google: Do the search "site: https://news.ycombinator.com "rackless ruth"" and find 0.

Indexing the entire web isn't instantaneous, you know.

Re: Google Memory Loss

#539
I wish you could use both Verbatim and Sort by date at the same time. Instead, you end up having to choose between recent but irrelevant results, or relevant but outdated results.

Re: Google Memory Loss

#540

Earlier quoted context omitted.

> has seemingly degraded into an approximation of a search engine that has knowledge of only very superficial information, will try to rewrite your queries and omit words (including the very word that makes all the difference...) I think the biggest irony is that the web allows for more adoption of long-tail movements than ever before, and Google has gotten significantly worse at turning these up. I assume this has s…

This same AI effect can be seen in the Android keyboard, where _properly_ spelled words will be replaced after typing another word or two because it's been determined to be more likely what you want. It's infuriating.

Its annoying autocorrection tendency to choose 'fir' instead of 'for' frustrates me. People almost never use the word 'fir', but use the word 'for' often. It would be nice if you could blacklist words you want it to never choose.

There is SwiftKey, on the hand, that does those kind of annoying corrections a couple of times, remembers your choice, and does them no more. It's been a long time since I've seen a 'fir' with SwiftKey.

Post reply on HN