When I was at Watson this is the first thing I told every customer: before you start with AI are you already doing the more mundane data science on your structured data? If not, you shouldn't go right away for the shiny object. This said I still believe the article is mistaken in its evaluation of potential impact (and its fuzzy metaphore of pipes). Unstructured or semi-structured or dirty data is much more prevalent…
And before you do mundane data science on your structured data, you should figure out if there is a better way to get cleaner raw data, more data, as well as more accurate data. For example, I predict stereo vision algorithms will die out soon, including deep-learning-assisted stereo vision. It's useful for now but not something to build a business around. Better time-of-flight depth cameras will be here soon enough.…
My experience is that waiting for cleaner data is often like waiting for Godot and will often be a project killer (sometimes justifiably). This is a key issue at the moment in advanced ML: clean large training sets ideal for supervised training are elusive and the companies making real-world advances are pretty much all using available data (and semi-supervised techniques) rather than expensive made up training sets.