Spark + MLlib, Python + Pandas + NumPy + Keras + TensorFlow + PyTorch, R, SQL, top placement in some Kaggle competitions. This would get you long way.
Good tool set recommendations (+1 for mentioning SQL, immensely helpful), and I enjoy Kaggle. Not sure how critical top placement is, though. It seems like getting into the upper echelons of Kaggle is a matter of refining your model, and I do wonder how much value these refinements offer over a more basic and general approach in a real world scenario. To be clear, when I say I wonder, I'm not saying I'm rejecting the…
Ask HN: As a data scientist, what should be in my toolkit in 2018?
121–130 of 177 posts
Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?
#122Earlier quoted context omitted.
> If you are a math undergrad how could I ever expect to know more math than you? Read through, and do all the exercises in, one textbook each for: 1. Calculus 2. Linear Algebra 3. Abstract Algebra 4. Analysis 5. Topology 6. Probability Theory 7. Number Theory ...more or less in that order. Make sure your calculus book covers single variable and multivariable calculus. Supplement with applied mathematical statistics.…
That doesn't seem like a good use of time. I've tried reading through and doing the exercises in an abstract algebra textbook. It's a lot of work and the applicability to real world problems is virtually non-existent. I think a more targeted approach would give you a better return on your time.
1. The context is knowing more math than someone who has an undergraduate degree in it,
2. Abstract algebra is part of such a degree, and contributes significantly to overall mathematical maturity, and
3. You can avoid some subjects in the short term, but in the long term you can’t progress further without a reasonable mastery of algebra and analysis.
Probability theory and linear algebra are heavily used in data science. You won’t be as competitive a candidate for a job if you don’t have a firm grasp of both subjects. At a certain point, linear algebra ceases to be distinct from abstract algebra, and those exercises you were doing become applicable to real world results.
Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?
#123Earlier quoted context omitted.
I have to admit this scares me just a little bit. I'm a senior sysadmin who is trying to lateral transition into data science, but I'm no math whiz, I'm just good at pragmatic use of tech stacks and have a generally analytical mind. If you are a math undergrad how could I ever expect to know more math than you? Of course a standard deviation should be easy, but your comment on math just stuck out to me.
> If you are a math undergrad how could I ever expect to know more math than you? Read through, and do all the exercises in, one textbook each for: 1. Calculus 2. Linear Algebra 3. Abstract Algebra 4. Analysis 5. Topology 6. Probability Theory 7. Number Theory ...more or less in that order. Make sure your calculus book covers single variable and multivariable calculus. Supplement with applied mathematical statistics.…
Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?
#124Earlier quoted context omitted.
Do people get careers as 'data scientists' without masters degrees? I'm heavily biases towards experts... but still would call it fair in the general case to call anyone doing data science without at least a masters degree or the equivalent mental toolkit more like a 'data quack'.
Unless you do a masters degree in Data Science, AI, CS, Statistics, or a related field, a masters degree only serves as proof that you are intelligent and work hard. Someone without a masters degree can still have those attributes, but they would just have to prove it some other way. For example, if you have a bachelors degree from a top engineering school (MIT, Cal Tech, Stanford, Berkeley, etc.) you have proven tha…
I think the commenter’s point is (implicitly) about specific, directly relevant Master’s degrees. Obviously a general Master’s wouldn’t provide much of an advantage. The difficulty isn’t demonstrating intelligence and work ethic, it’s demonstrating targeted expertise.
> People without a masters degree, but more business experience, bring a different perspective, and are often more business results focused, and potentially work more collaboratively than an individual who just graduated from a masters program.
To be honest with you, this sounds to me like complete speculation. I’m not saying it’s wrong; rather it seems like it’s at best unempirical, and at worst unfalsifiable. The qualifiers you’re using (like “potentially”, or “often”) don’t seem like strong heuristics.
I think it would be helpful to discuss straightforward job descriptions. For most real data science roles, I would not weight any of what you’ve listed (except collaboration) as being remotely as useful as demonstrable expertise in computer science and statistics. For candidates without a Master’s degree, I wouldn’t take business experience or lack thereof as a signal whatsoever - I’d look for a relevant heuristic to replace it.
Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?
#125Earlier quoted context omitted.
I have to admit this scares me just a little bit. I'm a senior sysadmin who is trying to lateral transition into data science, but I'm no math whiz, I'm just good at pragmatic use of tech stacks and have a generally analytical mind. If you are a math undergrad how could I ever expect to know more math than you? Of course a standard deviation should be easy, but your comment on math just stuck out to me.
> If you are a math undergrad how could I ever expect to know more math than you? Read through, and do all the exercises in, one textbook each for: 1. Calculus 2. Linear Algebra 3. Abstract Algebra 4. Analysis 5. Topology 6. Probability Theory 7. Number Theory ...more or less in that order. Make sure your calculus book covers single variable and multivariable calculus. Supplement with applied mathematical statistics.…
Skip abstract algebra, topology and analysis. If you find yourself in the same room as a number theory book, walk away slowly without making eye contact lest it cast a spell on you.
Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?
#126grep, cut, cat, tee, awk, sed, head, tail, g(un)zip, sort, uniq, split; curl; jq, python3
Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?
#127Earlier quoted context omitted.
I have to admit this scares me just a little bit. I'm a senior sysadmin who is trying to lateral transition into data science, but I'm no math whiz, I'm just good at pragmatic use of tech stacks and have a generally analytical mind. If you are a math undergrad how could I ever expect to know more math than you? Of course a standard deviation should be easy, but your comment on math just stuck out to me.
Maybe consider being a data engineer or a systems engineer? There's a pretty big demand for people that can set up, maintain, and assist the data scientists with the more complex tech stacks out there. In a former job as a systems engineer, I set up Hadoop clusters and helped manage data going into and out of it. And if you do decide to continue learning to become a data scientist, you'll already have a solid footing…
Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?
#128Earlier quoted context omitted.
> If you are a math undergrad how could I ever expect to know more math than you? Read through, and do all the exercises in, one textbook each for: 1. Calculus 2. Linear Algebra 3. Abstract Algebra 4. Analysis 5. Topology 6. Probability Theory 7. Number Theory ...more or less in that order. Make sure your calculus book covers single variable and multivariable calculus. Supplement with applied mathematical statistics.…
I would amend as follows: Skip abstract algebra, topology and analysis. If you find yourself in the same room as a number theory book, walk away slowly without making eye contact lest it cast a spell on you.
Sure, skip number theory. Like I said, you could swap that out.
Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?
#129Re: Ask HN: As a data scientist, what should be in my toolkit in 2018?
#130Earlier quoted context omitted.
I think data scientist, much like software engineer, is something you can call yourself without having any credentials whatsoever. It’s why technical interviews can be so brutal, unfortunately. There are a lot of frauds out there. Money attracts frauds. What’s the fizzbuzz test for data scientists anyway?
I'm a data engineer for a startup that's trying to hire its first data scientist. The range of candidates that apply with this title is massive. Defining our expectations has been challenging. My phone screen "fizzbuzz" is having them calculate a standard deviation from an array of data w/out with only basic operators (no numpy.std). Then explain why they choose population/sample and explain the difference. I studied…
I'm having a little trouble trying to parse that sentence. Could you explain it better?
Based on what I think is being asked, the question is essentially: What is a STD? I think this is a very straightforward and fair question.
For less Stat-y HNers: For normally distributed data, the STD is the root of the Variance. The Variance is just the average of the square of the difference between the data points to the mean. Essentially: Take a point, find the distance to the mean, square that, average over all points you've done that to. That's the variance. Root the variance, that's the STD.