Live data from Hacker News

Big Data's Big Problem: Little Talent

online.wsj.com

121–130 of 161 posts

Re: Big Data's Big Problem: Little Talent

#121
post #91

"claims of severe talent shortage in Big Data http://online.wsj.com/article/SB1000142405270230472330457736... Ok... where are the high salaries (500k$ a year)? No? No real shortage." https://twitter.com/#!/lemire/status/196245665951649793 Business has a shortage of "big data" folks in much the same way I have a "huge sailboat" shortage. Neither of us want to pay for it. We want it, but not for the going rate. Only on…

Even more,

I don't notice an effort to expand the workforce by training or by recruitment of non-traditional workers, etc.

The contra-logical statement "99% of programming applicants are unqualified" gets a lot of play in this field. But I would suggest something like "we can make 99% of applicants look like idiots with our circus-like hiring process".

Yes, we've decided we have a shortage once we decide on five arbitrary disqualifications, expect all applicants to work 18 hours a day and start yesterday having no time to get up to speed (so experience on earlier large systems, say, is indeed not useful).

Re: Big Data's Big Problem: Little Talent

#122
post #116
post #105

Earlier quoted context omitted.

pmb is (correctly) saying that there is, by definition, no shortage of big data people. There's just a shortage at the (obviously below market clearing) price employers wish to pay. Also, you're ignoring the steep lead time to become a deep expert in stats / ML -- most likely a PhD plus significant programming time plus work experience.

Of course, by that definition, there is never a shortage of anything.

By that definition, the only absolute shortage is in things that aren't available at any price.

There can still be a relative shortage. There's shortage of gold-per-pound relative to rocks-per-pound, for example

Re: Big Data's Big Problem: Little Talent

#123
post #109

Earlier quoted context omitted.

Who is paying $200k for entry besides maybe google?

Startups.* *Equity value, may vary unpredictably. And it's entry post-PhD, not entry from college.

So $120k and (very expensive) lottery tickets is what you're saying =P

Re: Big Data's Big Problem: Little Talent

#124
I've done this kind of thing most of my career, including doing it for NASA and Unilever Research. You can't really train an average graduate to do this. You need someone with a pretty highly developed integration between 1:intuitive/creative abilities, 2:mathematical/analytical skills, and 3:engineering/ability to make things happen. Add to that 4:work experience in the real world, and 5:ability to easily understand how things work in a field you delve into for the first time... And there's very few people in the world who can do this. At my previous work place we tried for a whole year to hire someone who would at have at least some of these skills and seems promising to develop the rest on the job. We couldn't find anyone although we interviewed about 30 different people (from about 500 resumes most of them with a PhD in ML from a good university). And this was in central London, UK.

Re: Big Data's Big Problem: Little Talent

#125
post #117

Earlier quoted context omitted.

Bags of money are already being waved around, that is not the problem. Wages are already moving north of $200k for these positions because you can't find people with the basic skills for any amount of money. Being a "data scientist" as currently defined in practice requires someone to be a polymath with skills that are individually high value and not commonly found together. Roughly speaking, you need some aptitude a…

Computational geometry??? That's a new one for me. Do you mean only linear/convex programming? Incidentally, I would really like to hear about the kind of Real Work that data scientists end up doing with TBs of data, because I'm always fuzzy on the details. MCMC? Variational methods? SVMs? Or is it more oriented towards frequentist statistical methods, applied at "web-scale"?

I mean actual computational geometry. Reality is significantly non-Euclidean in complicated ways that have to be accounted for if precision matters.

Spatio-temporal analytics or the processing of sensing data frequently requires this. For a simple example, the surface of the Earth is approximately an oblate spheroidal surface, not even a 2-sphere. You can use Euclidean approximations for many cartographic purposes but for analytics this can introduce large errors in the analysis. Understanding how to compute non-Euclidean geometry models is surprisingly useful.

Re: Big Data's Big Problem: Little Talent

#126

I've done this kind of thing most of my career, including doing it for NASA and Unilever Research. You can't really train an average graduate to do this. You need someone with a pretty highly developed integration between 1:intuitive/creative abilities, 2:mathematical/analytical skills, and 3:engineering/ability to make things happen. Add to that 4:work experience in the real world, and 5:ability to easily understand…

But you never tried training them. Sure, no one can do it right off the bat, without prior experience.

Re: Big Data's Big Problem: Little Talent

#127
post #91

"claims of severe talent shortage in Big Data http://online.wsj.com/article/SB1000142405270230472330457736... Ok... where are the high salaries (500k$ a year)? No? No real shortage." https://twitter.com/#!/lemire/status/196245665951649793 Business has a shortage of "big data" folks in much the same way I have a "huge sailboat" shortage. Neither of us want to pay for it. We want it, but not for the going rate. Only on…

Even more, I don't notice an effort to expand the workforce by training or by recruitment of non-traditional workers, etc. The contra-logical statement "99% of programming applicants are unqualified" gets a lot of play in this field. But I would suggest something like "we can make 99% of applicants look like idiots with our circus-like hiring process". Yes, we've decided we have a shortage once we decide on five arbi…

The base-level skill set is being a very good applied mathematician with some good computer science skills. This is why a lot of "data scientist" types have degrees in things like physics. A lot of the database ETL stuff can be learned.

This is the reason why I cannot be a "data scientist", despite being an expert in parallel algorithm design and with strong database ETL experience. It would require me spending a couple years studying mathematics in depth that I do not currently know. The vast majority of programmers are at least as deficient as I am in critical skills for these positions.

We train our data scientists at my company but we usually do not start with software engineers. Our feedstock is strong applied mathematicians with some programming skills because the mathematics part is by far the most difficult to train for someone who has not already been doing it for years.

Re: Big Data's Big Problem: Little Talent

#128
post #21

Earlier quoted context omitted.

How hard can it be, though? Like taking a normal CS person and making them versatile with hadoop and so on? Could it be done for 20K$?

How hard can it be? Very hard. You run into all types of candidates who just aren't there yet: people working on research that's irrelevant to real world applications, people who have done data analysis/BI work that brand themselves as "data scientists," those who have the pedigree but cannot process and explore real-world data, those who have good analytical chops but not the distributed or advanced modeling experie…

If it is that hard the bar is probably set too high. Most of the skills are learned on the job after all. Most smart PhDs who can program well and have sound knowledge of statistics can learn to do this stuff.

Re: Big Data's Big Problem: Little Talent

#129
post #116

Earlier quoted context omitted.

Of course, by that definition, there is never a shortage of anything.

This is a very good quersion, and made me think a little. Here's my stab at it: If we define the shortage as "shortage of people willing to do X for $200,000 a year", that's clearly a bad definition. You should just pay more (as earl suggested) to get what you want. But what if that's just not possible on a macro level? Consider if you have an aggregate demand of "the market needs a total of 500 Data scientists". If…

Well, I applaud your effort but you've missed things.

The market doesn't need anything, human beings need and desire things and, in the capitalist model, the market is the means to balance those needs and desires things.

If 50 entrepreneurs desire 500 data scientists for their enterprises and only 250 fifty such scientists are on offer, they'll bid up the prices until some of them decide "I don't want them that much" and then we're done.

Of course, if our 50 entrepreneurs all have vast, vast wealth and a very strong desire for those scientists, we may see them creating bootcamps for quick data scientist development or whatever. Then they might indeed wind-up another 250 data scientists and again, we're done with no mystery.

Of course, you could argue that markets aren't as efficient as some claim but that's just about irrelevant to our reasoning here since we're mostly reasoning by definitions and extreme and so the OP basically holds true - you say you want but you've shown you don't want it that much and so you're really blowing smoke...

Re: Big Data's Big Problem: Little Talent

#130
post #2

Actually that is silly -- McKensey should now that there is and will never be a talent shortage. There will only be shortage of talent at a particular wage rate. If the companies paid newly graduated 'data-scientists' (what other kind of scientists are there? The tea-leaf reading kind?) 200k/year then they would have a lot more. It is pretty simple economics.

Companies already do pay close to $200k/year for entry level data scientists. (what other kind of scientists are there? The tea-leaf reading kind?) "Data scientist" refers to the guy who can set up a hadoop cluster, do statistics on TBs worth of data, derive useful conclusions and speed it up by tweaking the low level data formats or microoptimizing the calculation. The issue is rarely paying these guys an extra $20k…

The problem as I see it is that most companies are looking for the all-in-one perfect candidate. There is indeed a shortage of such people.

Say you need someone who knows a lot about Hadoop and Amazon EC and is also intimately familiar with most learning algorithms and has a PhD. You are having trouble finding the guy. You start crying about "the big data talent shortage".

And here is the problem. Most PhDs have no experience with Hadoop or Amazon EC. Some of them might know Java well enough.

Now, consider a smart guy with PhDwho knows Java and has done something parallel with it, working on real "dirty" data. He can pick up Hadoop in no time from your software engineers. He will learn to tweak and optimize in his time - it is domain specific and cannot be learned off the job.

Will he be hired? Probably not. But people will keep crying about shortage.

Post reply on HN