Earlier quoted context omitted.
That's an absurd claim. Show me the translation benchmark where you saw such a result.
Does "every time I've challenged people about this on Hacker News" count as a benchmark? See e.g. https://news.ycombinator.com/item?id=35531558 .
Don't get me wrong. I'm sure that once fine tuned by a human for a specific language pair that such systems are better at performing literal translation. But the value proposition of deep learning here is that you don't need a large team of experts to laboriously train a given language pair, that the entire training process is largely unsupervised, and that the translations aren't hopelessly literal. The ML algorithms can pick up on idioms given a sufficiently large dataset.