Live data from Hacker News

What Data Can’t Do

newyorker.com

51–60 of 106 posts

Re: What Data Can’t Do

#51

The article seems to stop pretty early, as if something is missing. It’s an anecdote about a government incentive to have doctors see patients within 48 hours causing doctors to refuse scheduling patients later than 48 hours in order to get the incentive bonus. This is not an example of limits of data, but an example of perverse incentives.

Did you use reader mode? because I noticed that when I use firefox's reader mode it will cut off part of the article on the New Yorker site.

When I first opened the article in Firefox, I only saw the first two paragraphs (same as the OP). This was true whether it was in reader mode or not. I opened it in Chrome and saw the whole article.

I just tried opening it in Firefox now (a couple hours later) and I see the whole article. If I switch to reader mode I do see it's truncated about halfway through, but I think that's a separate issue from what the OP was seeing.

Re: What Data Can’t Do

#52
post #31
post #28

Earlier quoted context omitted.

i think this actually gets at what makes applied ML distinct from statistics as a practice, even though there is a ton of overlap. statisticians make assumptions 1 and 2, and think of themselves as trying to find the "correct" parameters of their model. people doing applied ML typically assume they don't know 1 (although they might implicitly make some weak assumptions like sub-gaussian to avoid fat tails, etc.) and…

Yes - this is pretty much exactly how I explain the difference between machine learning and statistics. Despite using similar models, the expertise required for 'doing statistics' (statistical inference) is actually very different from machine learning. Machine learning fits into the 'hacker mentality' well - try stuff out see what works. To do statistical inference effectively, you really do need to spend time learn…

But without some statistical knowledge, isn’t there a risk of a lack of understanding about the robustness of “what works”?

Re: What Data Can’t Do

#53
post #4

I am increasingly worried with people applying ML in everything without any rigour. Statical inference generally only works well in very specific conditions: 1 - You know the distribution of the phenomenon under study (or make an explicit assumption and assume the risk of being wrong) 2 - Using (1), you calculate how much data you need so you get an estimation error below x% Even though most ML models are essentially…

As currymj commented, this isn't accurate for ML, only for classical statistics. In ML (or more specifically deep learning), we make no distribution-based assumptions, other than the fundamental assumption that our training data is "distributed like" our test data. Thus, there aren't issues with fat-tailed distributions since we make no such normality assumptions. Indeed, with the use of autoencoders, we don't assume…

I agree that ML tends to put weaker assumptions on the data than classical statistics and that it's a good thing.

However most ML certainly makes distributional assumptions - they are just weaker. When you're learning a huge deep net with an L2 loss on a regression task, you have a parametric conditional gaussian distribution under the hood. It's not because it's overparametrized that there's no distributional assumption. Vanilla autoencoders are also working under a multivariate gaussian setup as well. Most classifiers are trained under a multinomial distribution assumption etc.

And fat-tailed distributions are definitely a thing. It's just less of a concern for the mainstream CV problems on which people apply DL.

Re: What Data Can’t Do

#54
post #38
post #34

This author has published a couple of articles like this at the New Yorker They all have this in common: the author works through some interesting and in some ways unusual cases where data or statistics have been improperly or naively applied, with some social costs. I really enjoy the articles themselves. Then the New Yorker packages it up with a cartoon and a headline and subheadline like "Big Data: When will it ea…

> *Numbers don’t lie, except when they do. Harford is right to say that statistics can be used to illuminate the world with clarity and precision. They can help remedy our human fallibilities. What’s easy to forget is that statistics can amplify these fallibilities, too. As Stone reminds us, “To count well, we need humility to know what can’t or shouldn’t be counted.”* I do have a problem with her conclusion here. Ar…

> Are numbers really lying if it's actually an incorrect data collection method or conflicting definitions of criteria for generation of certain numbers

Obviously it's a figurative metaphor, but it's pretty clearly a case of "this supposedly objective factual calculation is presenting an untruth."

Re: What Data Can’t Do

#55

Data is not a substitute for good judgment, for empathy, for proper incentives. The article focuses on governments and bureaucracies but there's no better example than "data-driven" tech companies, as we A/B test our key engagement metrics all the way to soulless products (with, of course, a little machine learning thrown in to juice the metrics). I wrote about this before: https://somehowmanage.com/2020/08/23/data-i…

I think it's almost worse in tech because it largely works. If the government sets a flawed metric, their real goal of pleasing their constituents has failed and theoretically they either have to fix it or lose political support.

But in tech, if your goal is just to make money, soullessly following data will often get you there, to the detriment of everyone else. Clickbait headlines will get you more views. Full-page popup ads will get you more ad clicks/newsletter subscriptions. Microtransactions will get you more sales. Gambling mechanics will get you more microtransactions.

You can say it's a flawed metric, but I think in the end, most people just actually care more about making money than they do about building a good product.

Re: What Data Can’t Do

#56
post #4

I am increasingly worried with people applying ML in everything without any rigour. Statical inference generally only works well in very specific conditions: 1 - You know the distribution of the phenomenon under study (or make an explicit assumption and assume the risk of being wrong) 2 - Using (1), you calculate how much data you need so you get an estimation error below x% Even though most ML models are essentially…

As currymj commented, this isn't accurate for ML, only for classical statistics. In ML (or more specifically deep learning), we make no distribution-based assumptions, other than the fundamental assumption that our training data is "distributed like" our test data. Thus, there aren't issues with fat-tailed distributions since we make no such normality assumptions. Indeed, with the use of autoencoders, we don't assume…

I dunno, there are definitely distribution-based assumptions—good luck working with skewed data. Most old-school techniques are kinda additive, so nobody's really been assuming a single distribution for practical applications.

Current ML techniques just work well for the kinds of problems people are applying them to, which is kind of a tautology. We should definitely seek to understand the theory behind stuff like dropout and not consider our lack of understanding a strength.

Re: What Data Can’t Do

#57

Data is not a substitute for good judgment, for empathy, for proper incentives. The article focuses on governments and bureaucracies but there's no better example than "data-driven" tech companies, as we A/B test our key engagement metrics all the way to soulless products (with, of course, a little machine learning thrown in to juice the metrics). I wrote about this before: https://somehowmanage.com/2020/08/23/data-i…

I've written the same sentence before! This is so cool! pardon the wall of text.

Here's my thesis, curious to hear your thoughts.

At some time around 2005, when efficient persistence and computation became cheap enough that any old f500-corp could afford to endlessly collect data forever, something happened.

Before 2005, if a company needed to make a big corporate decision, there was some data involved in making the decision, but it was obviously riddled with imperfections and aggregations biases.

Before 2005, executives needed to be seasoned by experience, to develop this thing we call "Good Judgement", that allows them to make productive inferences from a paucity of data. The Corporate Hierarchy was a contest for who could make the best inferences.

Post-2005, data collection is ubiquitous. Individuals and companies realized that you don't need to pay people with experience any more, you can simply collect better data, and outsource decision-making to interpretations of this data. The corporate hierarchy now is all about how can gather the "best" data, where "best" means grow the money pile by X% this quarter.

"Good Judgement" used to be expected from the CEO, down to at least 1-3 levels of middle management above the front-line people. Now, it appears (to me) to be mostly a feature of the C-Suite and Boards, and it's disappeared elsewhere. Long-term, high-performing companies seem to have a more diffused sense of good judgement. But these are rare. maybe they always have been?

Anyways, as we agree, this has a tendency to lead in problematic directions. Here's my thesis on "why".

Fundamentally, any "data" is reductive of human experience. It's like a photograph that captures a picture by excluding the rest of the world.

Few people seem to understand this analogy, because they think photographs are the ultimate record of an event. Lawyers understand this analogy. With the right framing, angle, lighting (and of course, with photoshop), you can make a photograph tell any story you want.

It's the same issue with data, arguably worse since we don't have a set of standard statistics. We have no GAAP-equivalent for data science (yet?).

Our predecessors understood that data was unreliable, and compensated for this fact by selecting for "Good Judgement". The modern mega-corps demonstrate that we don't have a good understanding of this today, evidenced by religious "data-driven" doctrine, as you describe.

People will say "hey! at least some data is better than no data!", to which I'll say data is useless and even harmful in lieu of capable interpreters. In 2021, have an abundance of data, but a paucity of people who are capable of critical interpretation thereof.

I don't know if it's a worse situation than we had 20 years ago. But it's definitely a different situation, that requires a new approach. I think people are taking notice of it, so I'm hopeful.

Re: What Data Can’t Do

#58
post #4

I am increasingly worried with people applying ML in everything without any rigour. Statical inference generally only works well in very specific conditions: 1 - You know the distribution of the phenomenon under study (or make an explicit assumption and assume the risk of being wrong) 2 - Using (1), you calculate how much data you need so you get an estimation error below x% Even though most ML models are essentially…

As currymj commented, this isn't accurate for ML, only for classical statistics. In ML (or more specifically deep learning), we make no distribution-based assumptions, other than the fundamental assumption that our training data is "distributed like" our test data. Thus, there aren't issues with fat-tailed distributions since we make no such normality assumptions. Indeed, with the use of autoencoders, we don't assume…

> In ML (or more specifically deep learning), we make no distribution-based assumptions, other than the fundamental assumption that our training data is "distributed like" our test data.

Okay, so that's about the same as classical statistics. You're just waiving the requirement to know what the distribution is. You are still assuming there exists a distribution and that it holds in the future when you apply the model. Sure you may not be trying to estimate parameters of a distribution, but it is still there and all standard statistical caveats still apply.

> Indeed, with the use of autoencoders, we don't assume a single distribution, but rather a stochastic process.

Classical statistics frequently makes use of multiple distrutions and stochastic processes.

Re: What Data Can’t Do

#59
This article is totally gibberish. It's a terrible mixture of many unrelated things. Just because those things all have something to do with data (anything can be presented in numeric form), it does not make their issues are about data.

First, the Tony Blair example is not about data. It is a failure of government planning. It's wrong politics and wrong economy.

The G.D.P. example is laughable. G.D.P. is never intended to be used to compare individual cases. What kind of nonsense is this?

And the IQ example. The results are backed by decades of extensive studies. The author thinks picking a few critics can invalidate the whole field. And look! The white supremacist who gave Asians the highest IQ, what a disgrace to his own ideology.

Many more. I feel it's kind of tactic to produce this kind of article. Just glue a bunch of stuff, throw together with somethings seem to be related, bam, you got an article.

Re: What Data Can’t Do

#60

This article is totally gibberish. It's a terrible mixture of many unrelated things. Just because those things all have something to do with data (anything can be presented in numeric form), it does not make their issues are about data. First, the Tony Blair example is not about data. It is a failure of government planning. It's wrong politics and wrong economy. The G.D.P. example is laughable. G.D.P. is never intend…

Gotta disagree here. This article is acknowledging a pattern, that data is misused in many different areas.

I think the problem goes even deeper, which is a misunderstanding of the scientific method. Good discussion about this topic here: https://news.ycombinator.com/item?id=26122712

Post reply on HN