Live data from Hacker News

Big Data's Big Problem: Little Talent

online.wsj.com

151–160 of 161 posts

Re: Big Data's Big Problem: Little Talent

#151

Earlier quoted context omitted.

Even more, I don't notice an effort to expand the workforce by training or by recruitment of non-traditional workers, etc. The contra-logical statement "99% of programming applicants are unqualified" gets a lot of play in this field. But I would suggest something like "we can make 99% of applicants look like idiots with our circus-like hiring process". Yes, we've decided we have a shortage once we decide on five arbi…

The base-level skill set is being a very good applied mathematician with some good computer science skills. This is why a lot of "data scientist" types have degrees in things like physics. A lot of the database ETL stuff can be learned. This is the reason why I cannot be a "data scientist", despite being an expert in parallel algorithm design and with strong database ETL experience. It would require me spending a cou…

This is the reason why I cannot be a "data scientist"

Are you worried about this outcome at all? Do you see yourself playing an important role on a data team, one with less modeling responsibilities but more infrastructure/DB responsibilities?

I'm considering this path and would love to hear your opinion.

Re: Big Data's Big Problem: Little Talent

#152
post #118
post #66

Earlier quoted context omitted.

This market is already (at least partially) covered by small consulting shops which provide sales front for competent freelancers who don't feel like doing the whole corporate networking&sales ritual.

Where do I find these small consulting shops?

Well, the companies I know tend to do their recruiting based on word-of-mouth recomendations - so they have to know you from some previous job or you need to be recommended by someone etc. It's a tiny sample of a couple companies though, and that probably doesn't generalize to all "small consulting shops".

Re: Big Data's Big Problem: Little Talent

#153

Earlier quoted context omitted.

I completely take most of your points, but I think that pretty much all quantitative PhD's are going to be close to "data scientists". Given that stats and explaining your research are requirements, all that's left is to train them to program, which a lot of people are already doing. As a matter of fact, since I heard about this big data stuff I've been honing my skills in this area, in case the hype actually manifes…

> I think that pretty much all quantitative PhD's are going to be close to "data scientists" Having taken all but one of the core requirements for a masters' degree in statistics at a university with a well-respected statistics department, I can tell you that's very much not true. The true challenges in data science have almost nothing to do with what you spend 90% of your time as a graduate student studying (whether…

Again, I completely see where you are coming from. However, the difference between a Masters and a PhD are huge, far bigger than the gaps between any other form of education.

In a PhD, essentially everything you learn is to master a particular topic, or solve some kind of problem. This can often involve programming (it did for me) and almost certainly involves statistics (again, it did for me). The most important characteristic of a PhD is that you learn all this yourself (I certainly did). For instance, I was the only person in my department to learn R (although there were some oldtime Fortran and C programmers in my department), and then I ended up learning some python and java along with bash to deal with data manipulation problems and administering psychological measures of the internet. These are the kinds of skills that lead into me possessing some of the skills needed to be a data scientist, and with some experience in the private sector, I'll get there.

Bear in mind that I (almost) have a Psychology PhD, and this would all have been far easier for me if I had worked in physics, chemistry or any of the harder sciences. So from my perspective, I can see that this is where the data scientists of the future are going to come from.

Note that I looked up the job market, and made a conscious decision to train myself in these kinds of skills throughout my PhD, but if you are not capable of performing this kind of analysis that you probably shouldn't be doing a PhD anyway.

I really don't see how programming upends all that grad students learn (though I would be delighted to hear your thoughts), as to me it just seemed like the application of logic with the aid of computers. I'm not that good a programmer though, certainly not outside the application of stats, but within the next few years I will be.

Re: Big Data's Big Problem: Little Talent

#154
post #49
post #40

Earlier quoted context omitted.

Whether or not you have these skills: potential employers need to SEE them. There are lots of pretenders out there, and employers are appropriately wary. Are you showing employers results on your web page that a worse ML practitioner can't match. Putting results from a kaggle competition on my LinkedIn page landed me my current job (and I am still contacted by potential employers every couple weeks). The employers of…

Did you win the competition ? If not what was your rank, if you don't mind sharing ? I'd be really interested to know how much impact this could have.

I was #4 at the time my current employer contacted me. I've continued in the competition with a small team of their employees, and we are currently #2.

I don't know how sensitive the number of job contacts is to your ranking. My hunch is that a lower ranking would still establish credibility.

Re: Big Data's Big Problem: Little Talent

#155

Earlier quoted context omitted.

The base-level skill set is being a very good applied mathematician with some good computer science skills. This is why a lot of "data scientist" types have degrees in things like physics. A lot of the database ETL stuff can be learned. This is the reason why I cannot be a "data scientist", despite being an expert in parallel algorithm design and with strong database ETL experience. It would require me spending a cou…

This is the reason why I cannot be a "data scientist" Are you worried about this outcome at all? Do you see yourself playing an important role on a data team, one with less modeling responsibilities but more infrastructure/DB responsibilities? I'm considering this path and would love to hear your opinion.

To be clear, I chose this outcome. I am good with mathematics but not the mathematics usually needed as a data scientist and I have relatively little interest in investing the time to learn. Being a data scientist is a great job for some people but probably not what I would choose even if I was a developer again.

There is a continuum of skill balances; some people are more "data" than "scientist" and vice versa. The most useful balance varies from job to job. There are plenty of opportunities for people that have strong skills standing up clusters even if you have relatively weak analysis and model building skills. I would not dissuade anyone from becoming a data scientist, it will pay very well for the foreseeable future, but the skill set requires real effort to acquire. At a small company there is likely opportunity to learn the trade by coming at it from the infrastructure side of things.

It is a young enough area that it should be pretty easy for talented individuals to invent a career if they apply themselves.

Re: Big Data's Big Problem: Little Talent

#156
post #83

Earlier quoted context omitted.

I agree somewhat, but retention is a fairly major problem. The shift away from career-length employment means that neither employers nor employees assume there will necessarily be much loyalty or longevity in the relationship. I think the decline in on-the-job training is directly related. Engineering firms used to be able to assume that it's okay to lose money on the first five years or so of an employee's work, if…

What prevented employees from jumping ship before? Was there better long term benefits associated with staying with a company?

Yeah, one of the long term benefits was that you could work there your whole career, if you wanted to. By the mid-80s though, it was clear that layoffs were to become a regular fixture of corporate life.

Re: Big Data's Big Problem: Little Talent

#157
post #91

"claims of severe talent shortage in Big Data http://online.wsj.com/article/SB1000142405270230472330457736... Ok... where are the high salaries (500k$ a year)? No? No real shortage." https://twitter.com/#!/lemire/status/196245665951649793 Business has a shortage of "big data" folks in much the same way I have a "huge sailboat" shortage. Neither of us want to pay for it. We want it, but not for the going rate. Only on…

The salaries are already moving north of $200k even outside of Silicon Valley and New York City and getting more expensive by the month. How high do they have to be before we have a "shortage"? The problem is not lack of money, it is that demand has greatly outstripped a finite supply. Very high wages do not automagically create new people with the requisite skills and this is the real bottleneck. It takes significan…

In context: $200k is roughly what attorneys are paid in their 5th-6th years at large firms that service large corporations. Lawyers are in surplus right now.

In finance, talented individuals are routinely paid well multiples of $200k for their work (even post-crash).

So while $200k is high for salaries generally, it certainly is not high enough to imply a shortage in a highly specialized field.

Re: Big Data's Big Problem: Little Talent

#158

Earlier quoted context omitted.

> I think that pretty much all quantitative PhD's are going to be close to "data scientists" Having taken all but one of the core requirements for a masters' degree in statistics at a university with a well-respected statistics department, I can tell you that's very much not true. The true challenges in data science have almost nothing to do with what you spend 90% of your time as a graduate student studying (whether…

Again, I completely see where you are coming from. However, the difference between a Masters and a PhD are huge, far bigger than the gaps between any other form of education. In a PhD, essentially everything you learn is to master a particular topic, or solve some kind of problem. This can often involve programming (it did for me) and almost certainly involves statistics (again, it did for me). The most important cha…

> In a PhD, essentially everything you learn is to master a particular topic, or solve some kind of problem. This can often involve programming (it did for me) and almost certainly involves statistics (again, it did for me).

Yes, but the original point was that more or less any quantitative PhD would be expected to have these skills.

> For instance, I was the only person in my department to learn R

Case in point - and I can tell you from knowing the PhD students that I if I found someone who knew how to program in any language not used primarily for statistical computation (R, Stata, Matlab, SAS etc.), I would consider them the exception, not the norm.

The exact opposite is true about a data scientist.

> but if you are not capable of performing this kind of analysis that you probably shouldn't be doing a PhD anyway.

Or you just don't care about those types of jobs - and apparently there are plenty of those, because many, if not the majority, of PhD students I can think of aren't looking for data science jobs.

> I really don't see how programming upends all that grad students learn (though I would be delighted to hear your thoughts), as to me it just seemed like the application of logic with the aid of computers.

It's not programming per se, but the computational power that it brings makes certain techniques feasible, and other concepts and methods aren't obsolete - just no longer optimal. This is really a comment about statistics specifically. Most job postings for data science positions mention some form of the phrase 'machine learning' - and if they don't, they often have that in mind. Unfortunately, while demand for machine learning dominates the job market, in the grand scheme of things, it's just one branch in the field of statistics, and its 'parent' branch was relatively obscure until very recently. To this day, if a PhD student finished their program having next to none of the required academic background for machine learning, I doubt most academics would bat an eyelid. It's just not considered important from an academic standpoint. It's unfortunate that we have such a disconnect between academic interest and industry demand, but it's very much the case.

A basic example that I often cite about how computational power has fundamentally changed statistics from how it was for the previous few decades is in our selection of estimators. (I often cite this because anybody who's ever taken a statistics class probably had this experience). In every introductory statistics class (and for many non-intro classes as well), when studying inference, you spend 90% of your time talking about estimators for which the first moment has an expectation of zero, and the 10% is a 'last resort' when no 'better' estimator exists. Who decided that the first moment was the most important? What about the second? Third?

Well, it turns out that the first moment is easier to calculate, and, by coincidence, it happens to be the most relevant when your dataset is small (say, between 30 and 100). But once you're talking about datasets with observations which number in the thousands (which is still 'small' by some standards today!), you'd be insane to throw out any estimator that converges at a linear rate (rather than at the rate of \sqrt{n}) just because it introduces a small bias.

But we do - and that's reflected in the the sheer amount of academic research and literature that discusses the former, and the sheer lack of that reflects the latter. In many cases, the theory exists, but it was developed in an era in which it could never feasibly be applied.

Vestiges of this era are visible even in many statistical software packages - for another basic example, regressions by default assume homoskedasticity in errors, even though this is almost never valid in real life. Why? Because in a previous era, everyone imposed this assumption, because while the theory behind the alternative had been developed, it was expensive to carry out in practice (it involves several extra matrix multiplications).

I'm painting with a broad brush, but the general picture still very much holds.

Re: Big Data's Big Problem: Little Talent

#159
post #43

Earlier quoted context omitted.

I see this in software dev consulting, and it's probably in many other fields as well too. I see companies looking for one person who's a highly-skilled DBA, sysadmin, developer and who can interact affably with all levels of people in a company, including customers, at the drop of a hat. Often it's because they had one 'magical' person who did all that, though usually not very well, and the next person (or team) who…

Perhaps slightly offtopic, but at least in Software dev consulting, there's a market for the multi-hatted individual in consulting directly, and pay is generally commensurate with the number of hats you can speak intelligently on to a customer. This isn't necessarily true everywhere -- when I lived in Memphis, TN, I often felt that I was dooming my professional career as, every time I ran into a challenge, I'd fork m…

It also depends on what it is you "go wide" on.

Early in my career I had the option of taking the proprietary track (mainframes, DEC/AXP, Microsoft Windows NT, several other proprietary platforms), or the open one (Unix, Linux, shell tools, GNU toolchain). One thing I realized was that just on an ability to get my hands on the tools I wanted to use, the open route was vastly more appealing. It's gotten somewhat better, but you could still easily pay $10-$20k just for annual licenses for tools you'd use.

The other benefit I found was that there was a philosophy of openness and sharing which permeated the open route as well. I've met, face to face, with the founders of major technological systems. And while there are many online support channels for proprietary systems, I've found the ones oriented around open technologies are more useful.

Dittos on training for proprietary systems. It's wonderful ... if you want to learn a button-pushing sequence for getting a task done, without a particularly deep understanding of the process. The skills I've picked up on my own or (very rarely) through training on open technologies have been vastly more durable.

Re: Big Data's Big Problem: Little Talent

#160

Earlier quoted context omitted.

Again, I completely see where you are coming from. However, the difference between a Masters and a PhD are huge, far bigger than the gaps between any other form of education. In a PhD, essentially everything you learn is to master a particular topic, or solve some kind of problem. This can often involve programming (it did for me) and almost certainly involves statistics (again, it did for me). The most important cha…

> In a PhD, essentially everything you learn is to master a particular topic, or solve some kind of problem. This can often involve programming (it did for me) and almost certainly involves statistics (again, it did for me). Yes, but the original point was that more or less any quantitative PhD would be expected to have these skills. > For instance, I was the only person in my department to learn R Case in point - an…

I completely see your point on estimators and the nature of many if not most statistics courses. I suppose that I was lucky enough to study non-parametric statistics in my first year of undergrad, as there's a subset within psychology that's very suspicious of all the assumptions required for the traditional estimators.

That being said, I think you're missing my major point which is that a PhD should be a journey of independent intellectual activity, so the courses one takes should be of little relevance, and so can therefore be downweighted in considering what PhD students actually learn. I accept that this is an idealistic viewpoint (FWIW, the best thing that ever happened to my PhD in this context was my supervisor taking maternity leave twice during my studies, which forced me to go to the wider world for more information about statistics).

I accept your point about machine learning not being a major focus of academia (well except for machine learning researchers), and I think its awful. Its very sad that the better methods for dealing with large scale data and complex relationships are only used by private companies. Its not that surprising, but it is sad.

That being said, I firmly believe that any halfway skilled quantitative PhD can understand machine learning, most of which is based on older statistical methods. It may not be taught (yet), but its not that much of a mindblowing experience. I do remember that when I first heard about cross-validation I got shivers down my spine, but that may just be a sad reflection on my interests rather than a more general point.

Post reply on HN