Live data from Hacker News

Goodbye, data science

ryxcommar.com

171–180 of 415 posts

Re: Goodbye, data science

#171
post #95

Earlier quoted context omitted.

>Meaning you could absolutely suck at your job or be incredible at it and you’d get nearly the same regards in either case. One of the things I don't like about statements like this said in a Data Science context, is that they are true outside of Data Science as well. Executives make big decisions, managers make smaller decisions, nobody can evaluate how good/bad they really were for months or years. Engineers build…

Not to get too off topic, but as a 35 year old engineer it seems the world in general has far fewer consequences than I was raised to expect. Everything from businesses with bullshit ideas flourishing at a loss, to January 6 even being possible (politics aside I expected the Capitol Police to crack a lot more skulls than they did once people started smashing windows), to the whole FTX situation and the tepid response…

I think this hits the nail on the head. People are learning the meritocracy they were taught growing up isn't real so why would they work 80 hour weeks for a 25% bonus instead of a 10-15% bonus. The calculus gets even worse when the bonus people get is insignificant to that of switching careers often.

For what you specifically experienced, my opinion, the bigger the organization the more inevitable this seems to become. To make things worse the size of the organization isn't limited to just a company or non-profit but to the size of all groups involved, i.e. a small charity or non-profit that's part of a huge government program is similar to a small engineering team in a huge tech company. They could do huge things or be completely worthless and so long as they pass along positive messages up the chain and the org or company as a whole is doing well then yay no consequences.

We're (hopefully) at the beginning of a cycle where companies realize they are causing apathy amongst the majority of the employed and hopefully experiment (and succeed) in providing meaningful pay raises to the lower echelons which will come at the short term costs of profits but are justified for long term productivity. Or we'll just keep divolving into a dystopia

Re: Goodbye, data science

#172

Calling it Data Science was a tell. Have you noticed how non-scientific things add "science" to the name to make it sound like it has scientific rigor? Data Science, Political Science, Social Science, Scientology Compare to Physics, Biology, Math.

I too am critical of data science, but I think that this is a bit unfair. The scientific part could be called just 'statistics'.

Should we call deep learning “statistics”? It’s mostly empirical. Should recommendation systems be called “statistics”? What about multi objective non convex optimization?

Not everyone is doing regression and classification all day.

The above were studied in the Computer Science curriculum at my university. I work with a statistician, like someone with a degree in mathematical statistics. They know nothing about any of those.

Re: Goodbye, data science

#173
post #6

> Nobody knew or even cared what the difference was between good and bad data science work. Meaning you could absolutely suck at your job or be incredible at it and you’d get nearly the same regards in either case. In my experience it's even a little bit worse than that. Approaches that are wrong from a statistics point of view are more likely to generate impressive seeming results. But the flaws are often subtle. A…

You should pay the price for data leakage very quickly in production.

Does management look at slides or AB test dashboards?

Re: Goodbye, data science

#174

Earlier quoted context omitted.

But if the page loads slowly or the UI is unresponsive, people notice. The output of Data Science is harder for non-specialists to evaluate.

For establishing competence, you still have to dig in to see what caused the slowness. A regular user can't tell you that.

> For establishing competence, you still have to dig in to see what caused the slowness.

Not as management. You just have to see that other people's similar sites are not slow with the same resources, therefore it is possible for your site not to be slow. You don't have to know why you're failing to know that the totality of the people you hired were not as good as the people those others hired.

This is of course barring management failure; but if you're failing at management, that's about the same as saying that your engineers were under-resourced.

Engineering competence is largely composed of the skills to figure out what is causing problems e.g. slowness. If you can't figure out what is causing the slowness, your engineers aren't good enough to figure out what is causing the slowness, qed.

That's different than data science.

Re: Goodbye, data science

#175
post #95
post #6

> Nobody knew or even cared what the difference was between good and bad data science work. Meaning you could absolutely suck at your job or be incredible at it and you’d get nearly the same regards in either case. In my experience it's even a little bit worse than that. Approaches that are wrong from a statistics point of view are more likely to generate impressive seeming results. But the flaws are often subtle. A…

>Meaning you could absolutely suck at your job or be incredible at it and you’d get nearly the same regards in either case. One of the things I don't like about statements like this said in a Data Science context, is that they are true outside of Data Science as well. Executives make big decisions, managers make smaller decisions, nobody can evaluate how good/bad they really were for months or years. Engineers build…

> One of the things I don't like about statements like this said in a Data Science context, is that they are true outside of Data Science as well. Executives make big decisions, managers make smaller decisions, nobody can evaluate how good/bad they really were for months or years. Engineers build something amazing, or build a house of cards, nobody cares as long as the money people are happy, even if the business use case turns out to be wrong in the long run.

This is purely anecdata, but I have found that this is more pronounced in a data science context. Managers and executives are (in my experience) more willing to admit they don't understand engineering work product and seek input from technical advisors, and executives and managers deal with decision making on a daily basis and understand that it can be nuanced. But since almost everyone reads financial reports or has to make a chart in Excel every now and then, they know enough to read someone else's analysis but not enough to recognize their knowledge gaps (particularly wrt advanced statistics).

Re: Goodbye, data science

#176
post #6

> Nobody knew or even cared what the difference was between good and bad data science work. Meaning you could absolutely suck at your job or be incredible at it and you’d get nearly the same regards in either case. In my experience it's even a little bit worse than that. Approaches that are wrong from a statistics point of view are more likely to generate impressive seeming results. But the flaws are often subtle. A…

> Approaches that are wrong from a statistics point of view

When OP talked about "the main bottleneck to my work" in terms of areas he would need to learn more about -- I was expecting him to talk about facility with statistical methods and using them appropriately!

I'm not sure what to take from the fact that he never did! I would like to ask him what he thinks about that!

Re: Goodbye, data science

#177
post #6

> Nobody knew or even cared what the difference was between good and bad data science work. Meaning you could absolutely suck at your job or be incredible at it and you’d get nearly the same regards in either case. In my experience it's even a little bit worse than that. Approaches that are wrong from a statistics point of view are more likely to generate impressive seeming results. But the flaws are often subtle. A…

The problem is that nobody actually wants data science. They want data pseudoscience.

And for the same reason that people tend to want pseudoscience instead of science in any other domain, too. Science is slow, tentative, and messy, and usually responds to questions with even more questions rather than with answers.

Pseudoscience tends to be much more concerned with exuding confidence and providing clean-cut answers. It's what happens when a desire for science meets a need for instant gratification. Along the way, things like blinding and controls and watching for bias and validating assumptions tend to get dropped when they're inconvenient or difficult to explain. And they're always inconvenient and difficult to explain.

Re: Goodbye, data science

#178
post #6

> Nobody knew or even cared what the difference was between good and bad data science work. Meaning you could absolutely suck at your job or be incredible at it and you’d get nearly the same regards in either case. In my experience it's even a little bit worse than that. Approaches that are wrong from a statistics point of view are more likely to generate impressive seeming results. But the flaws are often subtle. A…

I’ve always disliked how data science was positioned within companies as well, it’s outside the critical path of product and engineering, which means it becomes a mere abstraction to management (e.g. “throw that problem to the data science team and see what they come up with”), resulting in very vague and abstract requirements and, hence, deliverables. I think there is huge value in the discipline and technologies, but it unfairly gets relegated when not integrated to the whole product / engineering process. Hence, the title / concept of Data Engineer seems like a much better fit for this role within many companies.

Re: Goodbye, data science

#179
post #151

Earlier quoted context omitted.

yep, exact same feeling here. I had several years as a "data scientist" and it was a an almost totally bullshit job. the org bought into the hype and hired a cohort of us straight out of university, but then couldn't find anything data-science-y for us to actually do. what I actually ended up doing 95% of the time was taping together dodgy excel-based workflows using python scripts. it gave me a visceral appreciation…

Could you elaborate on what kind of "software engineering" you now do? For someone who also would like to get out of data science, mentions of "I became a software engineer" don't really help to clarify what kind of SWE is feasible for a data scientist with decent programming chops to get into.

I don't know how replicable my success is. but, depending on where you work, it may be that as a data scientist you can provide more value to your business purely using your software skills than with any kind of stats knowledge, by figuring out how to unfuck existing crufty bureaucratic workflows. this can be more directly useful than any amount of hyperparameter twiddling on some ridiculous neural network chimera. at my last job I could see so much of people's time wasted on fucking idiocy and my mind rebelled against it, I had this drive to rip it all out and Do It The Right Way(tm). and in doing that, I learned a lot about software development, tooling, version control, documentation, and so on. one thing led to another, and I had turned myself into a software engineer.

nowadays I do -- well, fudging slightly but you could describe it as "industrial automation control". writing libraries to provide convenient abstractions for controlling industrial equipment, writing robust scripts to drive that equipment, run physical tests on the $widgets we make, aggregate the experimental data, store it, etc. in the interviews they liked how (in my DS job) I had taken existing inefficient excel based workflows that had human-in-the-loop, and automated them, made unit tests, wrote docs, considered failure modes that nobody had considered before, things like that. and I just read about a fuckton of different stuff. for example in the interviews they wanted to know if I had worked with concurrency, I said I hadn't because it just didn't come up in the work I did. but I knew a little about it because I read voraciously, then I was able to answer all the theoretical questions they posed about locks and threads and async and so on. obviously that didn't mean I really knew about concurrency (that's a kind of deep metis that can only be acquired by practical experience and I'm still only scratching the surface of it), but it demonstrated that I had curiosity to learn about the field outside of the immediate things I worked on day to day.

during that job hunt I also had a strong offer from a company that wrote software for the visual effects industry and they wanted someone to improve their automated testing and continuous deployment frameworks. I didn't know much about CI but I knew about testing (pytest and hypothesis and things like that). they liked me talking about that kind of thing.

I guess the lesson is, if you are right now a data scientist and you want to be a software engineer, you can just decide to be that right now. be proactive and find a software problem to solve, and solve it. you don't have to ask permission to do this .. what are they going to do, tell you to stop being useful? note what you did, then figure out how to do the next thing better based on what you learned. your pay stub will say you're a data scientist, but you should just think of it as clandestine self-directed on-the-job training for your next job, so you can talk about it in the interviews. does that make sense?

Re: Goodbye, data science

#180
post #59

I have never understood the what a good ML engineer couldn't do and a Data scientist could in _majority_ situations. When you need a decision to be made based on data its just common sense risk analysis added together with basic statistics. I feel some good field training in statistics(Look up Andrew Gelman) a couple of good courses on Linear, Bayesian Regression is all you need, rest is just engineering skill. The d…

I think the qualifying term here is "good". I've worked with a surprising number of MLEs that don't really understand gradient descent or how most models really work under the hood. They certainly couldn't implement most things from scratch if they needed to (neither could most data scientists).

I used to think an MLE was a solid engineer who also had a strong quantitative and numerical computing background. The kind of engineer that always has a copy of Numerical Recipes handy, and if needed, could reimplement core components of statsmodels and sklearn in javascript.

I think after this current contraction in tech is over we'll see that most of the remaining "data scientists/MLEs" will be the type of engineer I imagine an MLE to be.

Post reply on HN