Live data from Hacker News

Machine Learning Is the New Statistics

danielmiessler.com

31–33 of 33 posts

Re: Machine Learning Is the New Statistics

#31
post #25

Earlier quoted context omitted.

> by determining the distribution of the process that the data was generated from. Well, each random variable has a distribution. And there are a few distributions that are common so are taught. Then, presto, bingo, too many students conclude that an important first step is to find a distribution. However, commonly in practice, with just samples and without more in mathematical assumptions, finding a distribution is…

>And made no more than meager, general assumptions about distributions. I don't see how assuming the data are IID from a Gaussian is a meager assumption.

> I don't see how assuming the data are IID from a Gaussian is a meager assumption.

It's not. Somehow we have failed to communicate accurately.

I wrote

> But, with just meager assumptions, commonly can still proceed and know that are still making a best L^2 approximation.

In that sentence, I didn't suggest that those "assumptions" were the Gaussian i.i.d. of the previous paragraph. Instead, I left the "assumptions" unspecified. Why? Because there is quite a variety available. But lots of the options are "meager".

Typically with more assumptions, can get more results. But for model fitting, building, constructing, discovering, whatever, can still get a lot with next to nothing in assumptions.

Re: Machine Learning Is the New Statistics

#32

Machine learning is subset of statistics. The standard text in ML, "The Elements of Statistical Learning" is authored by statistics Professors. Statistics is the new statistics. The rest is marketing bullshit.

I think names for fields are largely defined by social forces behind it not its ideas. Terms like "statistics" are sociological artifacts, not mathematical ones. The difference between stats and machine learning is analogous to say, the difference between Canada and the United States. They're neighbours, separated at birth, share many of the same values but still each have their own personality, worldview, goals and…

It's surprising to me that you would describe the field of statistics as a separate 'sociological artifact', but then refer to the actual definition of the term when using the abbreviated word 'stats', as in your sentence ' I have no stats to back up this claim'.

Statistics are tallied numbers and represent actual measured values. The field of statistics is concerned with tallied numbers collected, probability is concerned with the likelihood of those numbers being produced under specific assumptions, and machine learning is a process that uses statistics to verify and adjust the probability model being used for study.

Those are all definitions used by mathematicians and statisticians (who are a subset of mathematicians), not 'sociological artifacts'. Things don't sound like heresy to a statistician unless he is arguing implicit versus explicit logic. That is regardless of how it feels when he walk into his department.

There is no need to prescribe artifacts' if we can just keep the correct definitions clear and not conflate them.

Re: Machine Learning Is the New Statistics

#33

Machine learning is subset of statistics. The standard text in ML, "The Elements of Statistical Learning" is authored by statistics Professors. Statistics is the new statistics. The rest is marketing bullshit.

I finally figured out the best way to respond to this.

The central concept of Machine Learning is self-improvement of the models that are used, based on data alone.

That is NOT the central concept of Statistics.

So yes, Machine Learning may be related to, part of, semantically linked, a subset of, or whatever you want to say there, but the fact that Machine Learning is designed to self-improve is (in my opinion anyway) the reason it should be considered a "new" way of evaluating the world.

Saying it's all just Statistics sounds a lot like calling consciousness "just another information processing mechanism" (yawn). I know it's not that big of a difference, but it's similar.

Self-improvement of data analysis models matter enough to warrant the separate name and the attention that comes with it.

Post reply on HN