Live data from Hacker News

Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

super.gluebenchmark.com

51–60 of 246 posts

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#51

Earlier quoted context omitted.

The second one means "the enemy successfully got several of our aircrafts". Specifically, definition 3a or 3b for the verb form here: https://www.merriam-webster.com/dictionary/land So potentially the enemy captured the aircraft (3a) or destroyed them (3b).

If taking the "captured" interpretation, I think it could be reasonably inferred that they successfully landed the aircraft at an airfield afterwards (same meaning). This was my initial read of it and it does not seem strange to me on reflection. I would like also to point out that even if we do interpret the second as meaning "destroyed", the first could then be interpreted as a combat aviator shooting down an oppos…

Landed in the sense of a fisherman landing a marlin.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#52
I think classifying this as human level is misleading.

Look at the sub-scores on the page. One score that looks very different from humans is AX-b.

The SuperGlue paper provides more context about AX-b

https://arxiv.org/pdf/1905.00537.pdf

AX-b "is the broad-coverage diagnostic task, scored using Matthews’ correlation (MCC). "

This is how the paper describes this test

" Analyzing Linguistic and World Knowledge in Models GLUE includes an expert-constructed, diagnostic dataset that automatically tests models for a broad range of linguistic, commonsense, and world knowledge. Each example in this broad-coverage diagnostic is a sentence pair labeled with a three-way entailment relation (entailment, neutral, or contradiction) and tagged with labels that indicate the phenomena that characterize the relationship between the two sentences. Submissions to the GLUE leaderboard are required to include predictions from the submission’s MultiNLI classifier on the diagnostic dataset, and analyses of the results were shown alongside the main leaderboard. Since this broad-coverage diagnostic task has proved difficult for top models, we retain it in SuperGLUE. However, since MultiNLI is not part of SuperGLUE, we collapse contradiction and neutral into a single not_entailment label, and request that submissions include predictions on the resulting set from the model used for the RTE task. We collect non-expert annotations to estimate human performance, following the same procedure we use for the main benchmark tasks (Section 5.2). We estimate an accuracy of 88% and a Matthew’s correlation coefficient (MCC, the two-class variant of the R3 metric used in GLUE) of 0.77. "

If you look at the scores, humans are estimated to score 0.77. Google T5 scores -0.4 on the test.

How did T5 get such a high score if it scored so abysmally on the AX-b test?

The AX scores are not included in the total score.

From the paper: "The Avg column is the overall benchmarkscore on non-AX∗ tasks."

If the AX scores were included, the gap between humans and machines would be bigger than the current score indicates.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#53

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

I've never seen "landed" used as in the second sentence, but I was definitely able to understand from context that it was not being used to mean the same thing as in the first sentence.

You've never landed a fish?

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#54

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

From the abstract of the associated paper: "performance on the benchmark has recently surpassed the level of non-expert humans, suggesting limited headroom for further research." It occured to me that hn_throwaway_99's question, and the responses to it, is the sort of dialog in which one could find additional headroom for further research into natural language understanding. We can understand, for example, that while…

Limited headroom? Seems like they're assuming greater-than-human language ability is just impossible and will never be surpassed.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#55

Earlier quoted context omitted.

I've never seen "landed" used as in the second sentence, but I was definitely able to understand from context that it was not being used to mean the same thing as in the first sentence.

You've never landed a fish?

Have caught a fish though

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#56

Earlier quoted context omitted.

I've never seen "landed" used as in the second sentence, but I was definitely able to understand from context that it was not being used to mean the same thing as in the first sentence.

You've never landed a fish?

Is land an acceptable habitat for a fish?

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#57

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

The example directly below that: "Justify the margins" and "The end justifies the means" is the one I find dubious. Obviously the former could mean to format a document, but those exact words in that structure could be a demand for someone to justify a financial margin for example. It is both true and false depending on the context.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#58
One thing to always point out in these cases is that the human baseline isn't "how well people do at this task," like it's often hyped to be. It's "how well does a person quickly and repetitively doing this do, on average." The 'quickly and repetitively' part is important because we all make more boneheaded errors in this scenario. The 'on average' part is important because the errors the algo makes aren't just fewer than people, they're different. The algos often still get certain things wrong that humans almost never would.

This is really really super great, let's be clear. It's just not up to the hype "omg super human" usually gets.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#59

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

So many options for sentence number two.

- The enemy stole the aircrafts, and after some drama in flight managed to land several of them.

- The enemy used remote control to force them to land.

- The enemy used coercive force to force our pilots to land them.

- The enemy captured them.

- The enemy shot them down.

- During a friendly event while we set our differences with our enemy aside and agreed to fly each other's aircraft at an airshow for some reason, we landed several of theirs, and they landed several of ours.

- There was a hearing mistake and "energy" (as in energy beam beamed by a UFO) was accidentally transcribed as "enemy."

- The writer is just screwing with us.

- The writer is not a native speaker of English, and they made a mistake and actually meant that the enemy boarded several of our (parked) aircrafts.

- The writer is creative with language and believes that it would be cute to say that when an enemy projectile struck one of our aircrafts, then the enemy has "landed" that aircraft as one would land men on the moon or land rovers (no pun intended) on Mars.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#60

Earlier quoted context omitted.

I've worked in aviation for 8 years and also didn't understand this use of "landed". I've heard "grounded" used like this: "The maintenance issues gounded the jet," but not "landed".

I think the sentence is referring to aircraft that have been forced to land by the enemy, in contrast to "grounded" aircraft that had not taken flight. I haven't worked in aviation so my understanding of terminology could be wrong, but either way it is definitely an unusual example.

"The enemy landed 4 of our aircraft" without context wouldn't generally mean "forced to land" imo (as a native speaker). It would mean that they either destroyed them or managed to acquire them.

For example I might say that "they landed 4 aircraft with their daring" if they forced us to abandon an air craft carrier (e.g. by sinking it) and then managed to steal 4 of the planes (before it sunk). Or I might say "they landed 4 aircraft with that bomb" if they dropped a bomb on an airfield and it destroyed 4 aircraft.

Post reply on HN