Live data from Hacker News

Ask HN: As a data scientist, what should be in my toolkit in 2018?

news.ycombinator.com

121–130 of 177 posts

Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?

#121
post #119
post #52

Spark + MLlib, Python + Pandas + NumPy + Keras + TensorFlow + PyTorch, R, SQL, top placement in some Kaggle competitions. This would get you long way.

Good tool set recommendations (+1 for mentioning SQL, immensely helpful), and I enjoy Kaggle. Not sure how critical top placement is, though. It seems like getting into the upper echelons of Kaggle is a matter of refining your model, and I do wonder how much value these refinements offer over a more basic and general approach in a real world scenario. To be clear, when I say I wonder, I'm not saying I'm rejecting the…

Top placement in Kaggle attracts recruiters for higher positions; i.e. I observed a top 10 person getting a job of Head/VP of analytics in a large European company even if let's say formal education wasn't top 100. I agree real-world it is often useless, but people are drawn to proven winners.

Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?

#122
post #111

Earlier quoted context omitted.

> If you are a math undergrad how could I ever expect to know more math than you? Read through, and do all the exercises in, one textbook each for: 1. Calculus 2. Linear Algebra 3. Abstract Algebra 4. Analysis 5. Topology 6. Probability Theory 7. Number Theory ...more or less in that order. Make sure your calculus book covers single variable and multivariable calculus. Supplement with applied mathematical statistics.…

That doesn't seem like a good use of time. I've tried reading through and doing the exercises in an abstract algebra textbook. It's a lot of work and the applicability to real world problems is virtually non-existent. I think a more targeted approach would give you a better return on your time.

Sure, I agree. Abstract algebra isn’t directly helpful. But:

1. The context is knowing more math than someone who has an undergraduate degree in it,

2. Abstract algebra is part of such a degree, and contributes significantly to overall mathematical maturity, and

3. You can avoid some subjects in the short term, but in the long term you can’t progress further without a reasonable mastery of algebra and analysis.

Probability theory and linear algebra are heavily used in data science. You won’t be as competitive a candidate for a job if you don’t have a firm grasp of both subjects. At a certain point, linear algebra ceases to be distinct from abstract algebra, and those exercises you were doing become applicable to real world results.

Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?

#123
post #111

Earlier quoted context omitted.

I have to admit this scares me just a little bit. I'm a senior sysadmin who is trying to lateral transition into data science, but I'm no math whiz, I'm just good at pragmatic use of tech stacks and have a generally analytical mind. If you are a math undergrad how could I ever expect to know more math than you? Of course a standard deviation should be easy, but your comment on math just stuck out to me.

> If you are a math undergrad how could I ever expect to know more math than you? Read through, and do all the exercises in, one textbook each for: 1. Calculus 2. Linear Algebra 3. Abstract Algebra 4. Analysis 5. Topology 6. Probability Theory 7. Number Theory ...more or less in that order. Make sure your calculus book covers single variable and multivariable calculus. Supplement with applied mathematical statistics.…

Thank you for this list, going on the todo, along with every other relevant comment on this thread. (emacs org mode is my ds notebook and todo app)

Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?

#124
post #62

Earlier quoted context omitted.

Do people get careers as 'data scientists' without masters degrees? I'm heavily biases towards experts... but still would call it fair in the general case to call anyone doing data science without at least a masters degree or the equivalent mental toolkit more like a 'data quack'.

Unless you do a masters degree in Data Science, AI, CS, Statistics, or a related field, a masters degree only serves as proof that you are intelligent and work hard. Someone without a masters degree can still have those attributes, but they would just have to prove it some other way. For example, if you have a bachelors degree from a top engineering school (MIT, Cal Tech, Stanford, Berkeley, etc.) you have proven tha…

> Unless you do a masters degree in Data Science, AI, CS, Statistics, or a related field, a masters degree only serves as proof that you are intelligent and work hard. Someone without a masters degree can still have those attributes, but they would just have to prove it some other way.

I think the commenter’s point is (implicitly) about specific, directly relevant Master’s degrees. Obviously a general Master’s wouldn’t provide much of an advantage. The difficulty isn’t demonstrating intelligence and work ethic, it’s demonstrating targeted expertise.

> People without a masters degree, but more business experience, bring a different perspective, and are often more business results focused, and potentially work more collaboratively than an individual who just graduated from a masters program.

To be honest with you, this sounds to me like complete speculation. I’m not saying it’s wrong; rather it seems like it’s at best unempirical, and at worst unfalsifiable. The qualifiers you’re using (like “potentially”, or “often”) don’t seem like strong heuristics.

I think it would be helpful to discuss straightforward job descriptions. For most real data science roles, I would not weight any of what you’ve listed (except collaboration) as being remotely as useful as demonstrable expertise in computer science and statistics. For candidates without a Master’s degree, I wouldn’t take business experience or lack thereof as a signal whatsoever - I’d look for a relevant heuristic to replace it.

Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?

#125
post #111

Earlier quoted context omitted.

I have to admit this scares me just a little bit. I'm a senior sysadmin who is trying to lateral transition into data science, but I'm no math whiz, I'm just good at pragmatic use of tech stacks and have a generally analytical mind. If you are a math undergrad how could I ever expect to know more math than you? Of course a standard deviation should be easy, but your comment on math just stuck out to me.

> If you are a math undergrad how could I ever expect to know more math than you? Read through, and do all the exercises in, one textbook each for: 1. Calculus 2. Linear Algebra 3. Abstract Algebra 4. Analysis 5. Topology 6. Probability Theory 7. Number Theory ...more or less in that order. Make sure your calculus book covers single variable and multivariable calculus. Supplement with applied mathematical statistics.…

I would amend as follows:

Skip abstract algebra, topology and analysis. If you find yourself in the same room as a number theory book, walk away slowly without making eye contact lest it cast a spell on you.

Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?

#127

Earlier quoted context omitted.

I have to admit this scares me just a little bit. I'm a senior sysadmin who is trying to lateral transition into data science, but I'm no math whiz, I'm just good at pragmatic use of tech stacks and have a generally analytical mind. If you are a math undergrad how could I ever expect to know more math than you? Of course a standard deviation should be easy, but your comment on math just stuck out to me.

Maybe consider being a data engineer or a systems engineer? There's a pretty big demand for people that can set up, maintain, and assist the data scientists with the more complex tech stacks out there. In a former job as a systems engineer, I set up Hadoop clusters and helped manage data going into and out of it. And if you do decide to continue learning to become a data scientist, you'll already have a solid footing…

That might be an option to learn from, but it's not my end goal. As a senior sysadmin who was working for and reporting to PHD execs, I saw directly how what was needed was someone to do the data science and then bring convincing results and reports to the execs, essentially distilling the knowledge and wisdom of what needed to be done. I really want to fill that disconnect. (eg one of my failures as a sysadmin was me focusing too much on the technical, and now I want to expand and play the business board room politics game, but with data science)

Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?

#128
post #111

Earlier quoted context omitted.

> If you are a math undergrad how could I ever expect to know more math than you? Read through, and do all the exercises in, one textbook each for: 1. Calculus 2. Linear Algebra 3. Abstract Algebra 4. Analysis 5. Topology 6. Probability Theory 7. Number Theory ...more or less in that order. Make sure your calculus book covers single variable and multivariable calculus. Supplement with applied mathematical statistics.…

I would amend as follows: Skip abstract algebra, topology and analysis. If you find yourself in the same room as a number theory book, walk away slowly without making eye contact lest it cast a spell on you.

If you skip analysis and at least elementary topology your understanding will be limited to discrete probability, at best. If you skip abstract algebra, you’ll miss out on a lot of buildup to advanced linear transformations and operations in vector spaces.

Sure, skip number theory. Like I said, you could swap that out.

Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?

#130

Earlier quoted context omitted.

I think data scientist, much like software engineer, is something you can call yourself without having any credentials whatsoever. It’s why technical interviews can be so brutal, unfortunately. There are a lot of frauds out there. Money attracts frauds. What’s the fizzbuzz test for data scientists anyway?

I'm a data engineer for a startup that's trying to hire its first data scientist. The range of candidates that apply with this title is massive. Defining our expectations has been challenging. My phone screen "fizzbuzz" is having them calculate a standard deviation from an array of data w/out with only basic operators (no numpy.std). Then explain why they choose population/sample and explain the difference. I studied…

> ... of data w/out with only basic operators... (emphasis mine)

I'm having a little trouble trying to parse that sentence. Could you explain it better?

Based on what I think is being asked, the question is essentially: What is a STD? I think this is a very straightforward and fair question.

For less Stat-y HNers: For normally distributed data, the STD is the root of the Variance. The Variance is just the average of the square of the difference between the data points to the mean. Essentially: Take a point, find the distance to the mean, square that, average over all points you've done that to. That's the variance. Root the variance, that's the STD.

Post reply on HN