Live data from Hacker News

Machine learning’s crumbling foundations

pluralistic.net

61–70 of 106 posts

Re: Machine learning’s crumbling foundations

#61
post #50

It's a structural issue caused by the way wealth creation works for majority of people in tech. Job hopping, trendy frameworks in CV, "high-impact" projects done ASAP, etc. No one wants to do boring, slow pace work with lots of planning, reflection and introspection. And why would they do it? These kind of jobs are usually worst paid. We, the practitioners, have every economic incentive to go the other route. The pro…

Sadly PCR tests for COVID also test positive for flu and half a dozen other causes. That's why CDC/FDA are seeking proposals for a new test that actually works! https://www.cdc.gov/csels/dls/locs/2021/07-21-2021-lab-alert...

You've been repeating this, but it doesn't seem to be true...

https://news.ycombinator.com/item?id=28262833

Re: Machine learning’s crumbling foundations

#63
post #38
post #17

Earlier quoted context omitted.

Yeah, that was an odd claim about Han Chinese surnames. Many have been around for thousands of years ( https://www.chinadaily.com.cn/ezine/2007-07/20/content_54412... ) and almost all are single-character surnames based on a limited set of possible sounds (~400 in Mandarin, IIRC)

I don't know anything about Chinese surnames, but their paucity cannot be due to a limited set of possible sounds. First, I would interpret "sounds" as phonemes (including tones), and there are far fewer of those than 400. More likely what you mean is the number of combinations of phonemes into valid Chinese Mandarin monosyllables, of which I cannot imagine there being only 400. In any case, there are (from what litt…

You are getting something here by mentioning characters. Indeed there are a lot of distinct Chinese surnames written in the Chinese script that become identical after romanization, especially the romanization in the West where different tones are also ignored.

Wikipedia has a nice list of common Chinese surnames at https://en.wikipedia.org/wiki/List_of_common_Chinese_surname... and one can easily find examples: like 许 and 徐 both become Xu after romanization.

Re: Machine learning’s crumbling foundations

#64
post #50

Earlier quoted context omitted.

Sadly PCR tests for COVID also test positive for flu and half a dozen other causes. That's why CDC/FDA are seeking proposals for a new test that actually works! https://www.cdc.gov/csels/dls/locs/2021/07-21-2021-lab-alert...

You've fallen for the internet. Please restart and try again. https://www.reuters.com/article/factcheck-covid19-pcr-test-i...

Sigh. The deniers and antivaxers will have made up their minds already, and this will just be perceived as part of the mass media coverup. It's hopeless.

Re: Machine learning’s crumbling foundations

#65

It's a structural issue caused by the way wealth creation works for majority of people in tech. Job hopping, trendy frameworks in CV, "high-impact" projects done ASAP, etc. No one wants to do boring, slow pace work with lots of planning, reflection and introspection. And why would they do it? These kind of jobs are usually worst paid. We, the practitioners, have every economic incentive to go the other route. The pro…

I saw a large organization which was the epitome of this -- Executive Directors would propose ambitious ML projects, Directors would create plans and teams, Managers would execute on budgets, create more detailed plans, and then...someone actually needed to do the work.

Because of the length of the effort, the annual compensation would already have been handed out and the EDs, Directors, Managers had already "extracted" their compensation for the project, but usually had none left for the workers who eventually needed to do the actual work.

Not unexpectedly, a rough job was somehow jammed thru with understaffed, underpaid, and unmotivated low-level workers to actually "deliver" on the "AI" projects -- so victory could be declared at the top level...and new projects could begin.

This isnt an ML problem, i'm sure the whole cycle has been repeated with technology-of-the-day generation after generation. It has more to do with governance and organizational maturity to measure real impacts.

Re: Machine learning’s crumbling foundations

#66
post #56

> One common failure mode? Treating data that was known to be of poor quality as if it was reliable because good data was not available… they also use the poor quality data to assess the resulting models. This drives me nuts. Spend $10k getting high quality data and throw a simple model at it? Nah, let’s spend a month of time from someone making $400k/yr for less trustworthy results. And on the blogosphere it’s even…

I feel like its even worse than just resorting to bad data when good data isn't available: the field of deep learning has cultivated the perception that it's robust to bad data as one of its hallmarks. That is, you can pump relatively raw data into it and it will self-select features and then self-regulate their use so therefore most of the initial steps of data cleaning, feature selection etc are not necessary, or r…

The irony is that there are techniques for dealing with noisy/mislabeled/bad data (e.g. gold loss correction [0], errors-in-variables models [1]), but that stuff isn't "sexy" and not enough practitioners know about it.

0: https://arxiv.org/abs/1802.05300

1: https://en.m.wikipedia.org/wiki/Errors-in-variables_models

Re: Machine learning’s crumbling foundations

#67
post #30

Earlier quoted context omitted.

> Meanwhile, speech recognition seems to work extremely well by now (I am a little bit older, so I remember when it didn't work so well). *provided you speak English or Mandarin, the former preferably of a continental US variety It's astonishing how bad things get again once you mix in an accent, local dialect (e.g. Swiss German) or a less frequently spoken language (like Croatian).

Why would you expect dialects with vastly fewer training examples to be on par with the most widely spoken languages? It's a simple matter of available data, and the state of the art architectures operate on a paradigm that scales quality of the model to quantity of training data. If you want better speech recognition for Swiss-German, then record and transcribe hundreds of thousands of hours or whatever level of par…

> it's simply a function of the nature of these algorithms

Addendum: don't overlook the incentives and biases of the people building said algorithms.

Re: Machine learning’s crumbling foundations

#68

Earlier quoted context omitted.

I was tempted to just downvote this, but I thought I'd reply instead: No, they do not. An existing version of an app may get better over time, but unfortunately it then gets replaced with a different version, which starts from the position of extreme bugginess. In the case of Microsoft Office apps, for instance, one could easily argue that they are steadily getting worse as more and more features are added. Google Ch…

So why not go back to some old version of it? I don't think "memory consumption" is necessarily a good indicator, because sometimes using more memory is a sign of good optimization. Also how is the memory consumption if you turn off all modern features?

> So why not go back to some old version of it?

Because the old version doesn't work due to DRM/it depending on a remote API version that's no longer available/it's just flat out unavailable/etc...

> Also how is the memory consumption if you turn off all modern features?

It's cute you think you /can/ turn off the modern features in a lot of today's garbage.

Re: Machine learning’s crumbling foundations

#69
The article isn’t very clear about when harm has been done. It’s unclear which machine learning models researched production and whether human safety was on the line, like it would be for a bridge or a driverless car.

For example:

> Hundreds of ML teams built models to automate covid detection, and every single one was useless or worse.

That’s bad, but it doesn’t seem to mean there were hundreds that made it to production use? Drilling down, there is this bit from Technology Review [1]

> That hasn’t stopped some of these tools from being rushed into clinical practice. Wynants says it isn’t clear which ones are being used or how. Hospitals will sometimes say that they are using a tool only for research purposes, which makes it hard to assess how much doctors are relying on them. “There’s a lot of secrecy,” she says.

So clearly there was a lot of research that wasn’t immediately useful, but in the end it’s not clear how much reached production, whether it was critical to any health decisions, or whether people were harmed by it.

[1] https://www.technologyreview.com/2021/07/30/1030329/machine-...

Re: Machine learning’s crumbling foundations

#70
"Some of the data-cleaning workers are atomized pieceworkers, such as those who work for Amazon's Mechanical Turk, who lack both the context in which the data was gathered and the context for how it will be used."

The knowledge of the goal makes for yet another bias. "Need to know basis" in the intelligence "community" at the same time protects someone and tries to alleviate sources' biases.

Post reply on HN