Earlier quoted context omitted.
Modern machine learning depends heavily on data. If you have 100x the data, your algorithm is going to outperform a superior algorithm in most cases, because the data is the source of the intelligence. It's a big, scary barrier for startups, for innovation, and we're running the risk of creating data superpowers that can't be competed with. Some algorithms don't perform well until they have a threshold of data. When…
Actually the source of intelligence is not the data but the source of data. I've seen Norvig's video myself and while I see that they get better results with brute force on more data, that does not mean that better algorithms don't exist. In fact, I know for sure there are better algorithms (only because we have for example humans that are better than AI on certain tasks) that probably can use the available data in a…
If you're going to compete with Google as the one-stop-shop for searches, you're going to need to be able to provide search results that are consistently competitive. And I think Google has gotten good enough that you simply can't do that without access to some portion of their moutains of data. Nobody else has search history for every single American, and without that you are going to be crippled when {Random American} tries to search for something.