Live data from Hacker News

Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

super.gluebenchmark.com

81–90 of 246 posts

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#81
post #75

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

In looking through many of the replies to this downstream, it appears that the system is actually correct in that there's an obscure use of 'land' at play in the second sentence. It makes me think that there's going to be many adversarial examples of text that humans parse one way because of common usage while machines parse another way because of details like this.

Colorless green ideas sleep furiously!

Search for it if you’re interested in its origin.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#82
They came up with the SuperGLUE benchmark because they found that the GLUE benchmark was flawed and too easy to game. There were correlations in the dataset that made it possible to get questions right without real understanding, and so the results didn't generalize.

Could the same thing happen again with the better benchmark due to more subtle correlations? These things are tough to judge, so I'd say wait and see if it turns out to be a real result.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#83
I attended one of the talks(1) of the Sam Bowman. His talk was about "Task-Independent Language Understanding" and he also talked about GLUE and super GLUE; he mentioned that some models are passing an average person in experiments. They did some experiments to understand BERT's performance (2). (similar to article 'NLP's Clever Hans Moment') But they found a different answer to question "what BERT really knows," so he was skeptical about all conclusions. Check these out if you are interested in.

(1)[https://www.nyu.edu/projects/bowman/TILU-talk-19-09.pdf]

(2)[https://arxiv.org/abs/1905.06316]

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#84

Earlier quoted context omitted.

One of my favorite examples that I heard in a David Rock talk which I can no longer find on youtube: "Time flies like an arrow": Time moves swiftly and in one direction. Record the speed of flies in the same way you would an arrow. Time flies, which are a kind of fly, are fond of an arrow. (e.g. Time flies like an arrow, fruit flies like a banana).

"I eat my rice with butter." could mean that you use butter as a utensil to eat your rice with. There is often an unlikely way of parsing the sentence that gives an alternate meaning. The point is to test the computer to see if it can distinguish the likely parse from an unlikely one.

These aren't really alternate _parses_ though (in the sense that they don't give different parse trees). They do highlight the different possible meanings of "with" though.

I think "I eat my rice with chicken" vs "I eat my rice with children" vs "I eat my rice with chopsticks" is the canonical example here.

There's a whole field in NLP involved in showing what changes happen to entities mentioned in a sentence as a a side effect of the sentence, and this example shows it pretty well.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#85
Assuming that the baseline human score was set according to the performance of adult humans, then according to these results T5 has a language understanding ability at least as accurate as a human child.

In fact it's not just T5 that should be able to understand language as well as a human child, but also BERT++, BERT-mtl and RoBERTa, each of which has a score of 70 or more. There really shouldn't be anything else on the planet that has 70% of human language understanding, other than humans.

So if the benchmarks mean what they think they mean, there are currently fully-fledged strongly artificially intelligent systems. That must mean that, in a very short time we should see strong evidence of having created human-like intelligence.

Because make no mistake: language understanding is not like image recognition, say, or speech processing. Understanding anything is an AI-complete task, to use a colloquial term.

Let's wait and see then. It shouldn't take more than five or six years to figure out what all this means.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#86

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

I'm not a native English speaker and it is pretty obvious to me what both sentences mean.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#87
My experience with image classification benchmarks was that they approached human levels only because the scoring only counts how much they get “right” and doesn’t penalize completely whack answers as much as they should (like getting full credit for being pretty sure a picture of a dog was either a dog or an alligator). I suspect there’s something similar going on in these language benchmarks.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#88

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

The example looks like they're not written by native English speaker. It's funny reading English tests from other countries that are not English speaking because a lot of it focus on pedantics that are long lost while following a convention that would to us just feel _different_.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#89

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

I've never seen "landed" used as in the second sentence, but I was definitely able to understand from context that it was not being used to mean the same thing as in the first sentence.

Have you ever "landed" a deal? Or "landed" first strike in a game?

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#90

Earlier quoted context omitted.

The example directly below that: "Justify the margins" and "The end justifies the means" is the one I find dubious. Obviously the former could mean to format a document, but those exact words in that structure could be a demand for someone to justify a financial margin for example. It is both true and false depending on the context.

One of my favorite examples that I heard in a David Rock talk which I can no longer find on youtube: "Time flies like an arrow": Time moves swiftly and in one direction. Record the speed of flies in the same way you would an arrow. Time flies, which are a kind of fly, are fond of an arrow. (e.g. Time flies like an arrow, fruit flies like a banana).

It sounds like you're talking about garden-path sentences [0], and in particular: "time flies like an arrow; fruit flies like a banana" [1]. These are sentences whose structure tricks the reader into making an incorrect parse. My favourite of these has always been: "The horse raced past the barn fell".

[0] https://en.wikipedia.org/wiki/Garden-path_sentence

[1] https://en.wikipedia.org/wiki/Time_flies_like_an_arrow;_frui...

Post reply on HN