Live data from Hacker News

Machine learning’s crumbling foundations

pluralistic.net

101–106 of 106 posts

Re: Machine learning’s crumbling foundations

#101

Earlier quoted context omitted.

> So why not go back to some old version of it? Because the old version doesn't work due to DRM/it depending on a remote API version that's no longer available/it's just flat out unavailable/etc... > Also how is the memory consumption if you turn off all modern features? It's cute you think you /can/ turn off the modern features in a lot of today's garbage.

Pretty sure you can turn off a lot of things in modern browsers, if you find the hidden settings menu. For sure you can turn off things like JavaScript or video.

Pro tip: In fact, turning off JavaScript on a specific site is often a good way to get past their paywall.

On other sites, it just makes it not work. YMMV.

Re: Machine learning’s crumbling foundations

#102

> Ethnic groups whose surnames were assigned in recent history for tax-collection purposes (Ashkenazi Jews, Han Chinese, Koreans, etc) have a relatively small pool of surnames and a slightly larger pool of first names. This is... not accurate. The reason the Chinese have a small pool of surnames is that their surnames are much less recent than ours, not more recent. And I don't think the Ashkenazi surnames are partic…

Most Chinese surnames are ancient, but their romanizations are not:

* Names with the same pronunciation (homonyms) will probably end up with the same romanization. No romanization system has ever completely solved this problem, except by assigning arbitrary spelling differences.

* People will probably tell authorities the pronunciation of their name in their native language (Mandarin/northern dialects, Cantonese, Hakka etc.), which creates different-sounding versions of the same name. Worse, these romanizations are far from unified and don't correspond to the standardized romanization systems. Familiar example: Lee for the name 李 (Pinyin: Lǐ) and its homonyms.

* There are multiple romanization systems in use, which also yield different versions of a name, even if they all sound the same. A familiar example is Mao Zedong, whose name is Mao Tse-tung according to the Wade-Giles system, from before Pinyin romanizations of names became commonplace. For Cantonese names, a dizzying amount of romanization systems exist.

All of these render any surveys about Chinese names in the Diaspora extremely difficult, and most statistics completely garbage. For some families it might be impossible to recover the actual surname.

Re: Machine learning’s crumbling foundations

#103
post #102

> Ethnic groups whose surnames were assigned in recent history for tax-collection purposes (Ashkenazi Jews, Han Chinese, Koreans, etc) have a relatively small pool of surnames and a slightly larger pool of first names. This is... not accurate. The reason the Chinese have a small pool of surnames is that their surnames are much less recent than ours, not more recent. And I don't think the Ashkenazi surnames are partic…

Most Chinese surnames are ancient, but their romanizations are not: * Names with the same pronunciation (homonyms) will probably end up with the same romanization. No romanization system has ever completely solved this problem, except by assigning arbitrary spelling differences. * People will probably tell authorities the pronunciation of their name in their native language (Mandarin/northern dialects, Cantonese, Hak…

None of that is relevant here - every effect you list (well, not the first one) tends to increase the perceived variety of Chinese surnames, while the observation we're explaining is a lack of variety. That lack of observed variety is due to an actual lack of variety which your effects have failed to mask. And that actual lack of variety is due to the age of the system.

Re: Machine learning’s crumbling foundations

#104

Earlier quoted context omitted.

I saw a large organization which was the epitome of this -- Executive Directors would propose ambitious ML projects, Directors would create plans and teams, Managers would execute on budgets, create more detailed plans, and then...someone actually needed to do the work. Because of the length of the effort, the annual compensation would already have been handed out and the EDs, Directors, Managers had already "extract…

That sounds truly awful. Not necessarily surprising —- but could you give us some clues as to which company this was so that we can avoid working there?

This is a common enough occurrence that you probably want to learn how to spot bad setups, where-ever they might be. I think the key is to discern Value vs Vanity. You want to be on value-add projects (those producing revenue, or reducing risk, increasing speed, or reducing cost) but not on Vanity projects.

The trouble is that differentiating Vanity vs Innovation is hard. You can discern them though, in two ways I think:

1. By the level of motivation of low-level workers (true innovation is exciting) while underfunded vanity projects are soul crushing

2. By the seeming intentions of senior management -- are they more focused on the stated goal or on press/buzz?

I do not have sufficient n-value to come up with hard and fast rules but i'd love to hear others' thoughts

Re: Machine learning’s crumbling foundations

#105
post #102

Earlier quoted context omitted.

Most Chinese surnames are ancient, but their romanizations are not: * Names with the same pronunciation (homonyms) will probably end up with the same romanization. No romanization system has ever completely solved this problem, except by assigning arbitrary spelling differences. * People will probably tell authorities the pronunciation of their name in their native language (Mandarin/northern dialects, Cantonese, Hak…

None of that is relevant here - every effect you list (well, not the first one) tends to increase the perceived variety of Chinese surnames, while the observation we're explaining is a lack of variety. That lack of observed variety is due to an actual lack of variety which your effects have failed to mask. And that actual lack of variety is due to the age of the system.

If your point is the age: most of these names are ancient and can be traced to earlier that the first millenium BC. This is long enough that names can actually start to die out. Also, family names carry great significance in East Asian cultures*. They can carry great prestige, but also infamy. People often changed family names to become less associated with disgraced people. Emperors awarded their surname to loyal and meritous commoners, and these in turn gave it to their followers. Sometimes, whole populations adopted them. This happened to the Li (李), Chen (陳) and Wang (王) surnames.

It's true though that there are actually not that many to begin with. It's just a quite restricted set of words, and because of the writing system there is no variety because of spelling differences. Most surnames are only one character long, and the really long ones are mostly transliteration of non-Han surnames. Also, many non-Han populations were assigned a common surname when they became sinicized.

*: Western family names are mostly rooted in patronymics, place names, professions or adjectives.

Re: Machine learning’s crumbling foundations

#106
post #63
post #38

Earlier quoted context omitted.

I don't know anything about Chinese surnames, but their paucity cannot be due to a limited set of possible sounds. First, I would interpret "sounds" as phonemes (including tones), and there are far fewer of those than 400. More likely what you mean is the number of combinations of phonemes into valid Chinese Mandarin monosyllables, of which I cannot imagine there being only 400. In any case, there are (from what litt…

You are getting something here by mentioning characters. Indeed there are a lot of distinct Chinese surnames written in the Chinese script that become identical after romanization, especially the romanization in the West where different tones are also ignored. Wikipedia has a nice list of common Chinese surnames at https://en.wikipedia.org/wiki/List_of_common_Chinese_surname... and one can easily find examples: like…

This effect is reduced by the fact that there are multiple romanization systems in use.

Even taking that into account, the diversity is quite low. There are a lot of Chinese characters, but very few of them are actually used as surnames. Less than ~500 names are shared by more than 95% of the population, and most current surveys arrive at way less than 10000 in active use. There are two- and three-character names, but because of their low occurrence they are even more vulnerable to various processes that reduce the diversity of family names.

Chinese family names carry great significance. They were often assigned as rewards, and people adopted different ones for various reasons. This makes their distribution and future development subject to more than random chance.

Post reply on HN