Earlier quoted context omitted.
Well, two things there - First, the case you describe is one that's covered by using Spark for batch processing. Parent was criticizing using it and other tools for stream processing. The two are very different use cases. Second, just gotta call out that 2nd paragraph. A qualified data scientist should have a solid training in statistics. And someone who has a solid training in statistics should rarely if ever make t…
Right, for batch models and training evaluation. Re. assumptions -- it all depends on the data. If your data, say, collection of bird songs on Galapagos, then assuming there's correlation between day-of-week is counter productive. If a data scientist says me that they need to spend a day or two checking that assumption is correct, then I would look for a more productive scientist. Time to market is critical. And ther…
The counter-hypothetical you're giving just isn't plausible for the situation at hand. Analyzing bird songs on the Galapagos is scientist work, not data scientist work. Analyzing bird songs in an urban area might more plausibly be data scientist work, in that it gets you back toward business applications, but also gets us back into a spot where there's no way you could just assume there's no day-of-week effect.
I'm gonna stand firm on that one. That sort of thing is worrisome - I don't care if it's someone who holds a PhD or someone who just took a MOOC or two, data scientists have a professional responsibility to be way less sloppy than that.