Live data from Hacker News

A Convolutional Neural Network for Modelling Sentences (2014) [pdf]

arxiv.org

21–25 of 25 posts

Re: A Convolutional Neural Network for Modelling Sentences (2014) [pdf]

#21
post #19
post #10

Earlier quoted context omitted.

Notably, Mikolev is now at Facebook. My hunch (as a total outsider) is that anything Google publishes is about 2 years behind their current best practices.

The back-and-forth on the ImageNet record between Google/Facebook/Baidu suggests that, unless they're exquisitely coordinating research release, at least some of what they're releasing is indeed close to their state of the art . obviously writing it up does mean the specified results will be a bit behind what they can actually do in their labs, but that's true of everywhere

Not sure why that follows. (Can you outline the reasoning a bit more?)

As an example model, imagine all three are years ahead of published material, and each has policies that encourage publishing only the minimum necessary to "take the crown" (because any more would risk diluting proprietary advantages).

That'd result in the observed back-and-forth, and – given lead-times on paper-writing, internal review, and marquee conferences – it wouldn't necessarily force an acceleration of the "published state-of-the-art" up to the level of all of their "internal states-of-the-art". (Perhaps, for a group in trailing/catch-up position, their published results will be very close to their best. But the leader could be arbitrarily further ahead.)

By analogy to English auction bidding: outsiders only learn the second-highest reserve price just before the end, and never learn the winner's true reserve price. But here there's no end, and there's always potential competitive reasons for the top-N pack to hide some of their leading-edge practices.

This doesn't require explicit coordination… but there could be explicit coordination, too! All the programs are staffed by former students/colleagues/coworkers of each other.

Re: A Convolutional Neural Network for Modelling Sentences (2014) [pdf]

#22
post #21
post #19

Earlier quoted context omitted.

The back-and-forth on the ImageNet record between Google/Facebook/Baidu suggests that, unless they're exquisitely coordinating research release, at least some of what they're releasing is indeed close to their state of the art . obviously writing it up does mean the specified results will be a bit behind what they can actually do in their labs, but that's true of everywhere

Not sure why that follows. (Can you outline the reasoning a bit more?) As an example model, imagine all three are years ahead of published material, and each has policies that encourage publishing only the minimum necessary to "take the crown" (because any more would risk diluting proprietary advantages). That'd result in the observed back-and-forth, and – given lead-times on paper-writing, internal review, and marqu…

> That'd result in the observed back-and-forth, and – given lead-times on paper-writing, internal review, and marquee conferences – it wouldn't necessarily force an acceleration of the "published state-of-the-art" up to the level of all of their "internal states-of-the-art".

If they aren't coordinating very carefully, a single defector would result in the entire slack being used up by a single announcement; and the bigger the slack used up, the more a PR win it is...

Re: A Convolutional Neural Network for Modelling Sentences (2014) [pdf]

#23
post #22
post #21

Earlier quoted context omitted.

Not sure why that follows. (Can you outline the reasoning a bit more?) As an example model, imagine all three are years ahead of published material, and each has policies that encourage publishing only the minimum necessary to "take the crown" (because any more would risk diluting proprietary advantages). That'd result in the observed back-and-forth, and – given lead-times on paper-writing, internal review, and marqu…

> That'd result in the observed back-and-forth, and – given lead-times on paper-writing, internal review, and marquee conferences – it wouldn't necessarily force an acceleration of the "published state-of-the-art" up to the level of all of their "internal states-of-the-art". If they aren't coordinating very carefully, a single defector would result in the entire slack being used up by a single announcement; and the b…

OK, but in this sort of race, publishing your very latest internal techniques lets everyone else instantly catch up.

If these techniques are commercially valuable and in current use – and I believe they are! – then no leading company, or even member of the leading pack, would want to do a current-best reveal. They all have competitive reasons to do only carefully vetted, incremental reveals of somewhat-older work.

Re: A Convolutional Neural Network for Modelling Sentences (2014) [pdf]

#24
post #23
post #22

Earlier quoted context omitted.

> That'd result in the observed back-and-forth, and – given lead-times on paper-writing, internal review, and marquee conferences – it wouldn't necessarily force an acceleration of the "published state-of-the-art" up to the level of all of their "internal states-of-the-art". If they aren't coordinating very carefully, a single defector would result in the entire slack being used up by a single announcement; and the b…

OK, but in this sort of race, publishing your very latest internal techniques lets everyone else instantly catch up. If these techniques are commercially valuable and in current use – and I believe they are! – then no leading company, or even member of the leading pack, would want to do a current-best reveal. They all have competitive reasons to do only carefully vetted, incremental reveals of somewhat-older work.

You can publish papers without giving away all the details necessary to make it work. Sad but true.

For example, submitted to HN sometime ago was the blog of a dude trying to make Deepmind's 'neural turing machine' work; he was having a hella time because the published paper seems to be unclear or skip over a number of crucial points. Or more historically, the German chemical giants made an art of filing patents on all their key techniques, to gain IP protection, but leaving out enough crucial details that when the USA gleefully seized their IP rights during WWI, the American companies discovered they couldn't make the processes work.

Re: A Convolutional Neural Network for Modelling Sentences (2014) [pdf]

#25
post #24
post #23

Earlier quoted context omitted.

OK, but in this sort of race, publishing your very latest internal techniques lets everyone else instantly catch up. If these techniques are commercially valuable and in current use – and I believe they are! – then no leading company, or even member of the leading pack, would want to do a current-best reveal. They all have competitive reasons to do only carefully vetted, incremental reveals of somewhat-older work.

You can publish papers without giving away all the details necessary to make it work. Sad but true. For example, submitted to HN sometime ago was the blog of a dude trying to make Deepmind's 'neural turing machine' work; he was having a hella time because the published paper seems to be unclear or skip over a number of crucial points. Or more historically, the German chemical giants made an art of filing patents on a…

Indeed, and if that's also happening with the deep-learning papers here, then it's another mechanism in support of my main point: what's published is often, by motivated choice, behind what's being done internally. (And not just, "time it takes to write up" behind.)

Given that intent, there's no way to deduce from the pattern of new claimed results whether what's being revealed is a merely a few months, or many years, behind.

Post reply on HN