Live data from Hacker News

Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

super.gluebenchmark.com

41–50 of 246 posts

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#41
post #16

There was an article[1] posted to HN recently about these benchmarks, and it was pretty skeptical. Regarding SuperGLUE specifically, it asked: "Indeed, Bowman and his collaborators recently introduced a test called SuperGLUE that's specifically designed to be hard for BERT-based systems. So far, no neural network can beat human performance on it. But even if (or when) it happens, does it mean that machines can really…

This feels hollow. Can't this be said about any benchmark? It seems natural and proper that as one benchmark becomes saturated, we introduce harder benchmarks. I don't think anyone in the field thinks that once we match human performance on benchmark X, we're officially done. It just means it's time for more interesting benchmarks. Over time, if it starts to become difficult to design benchmarks that humans can outpe…

> Can't this be said about any benchmark?

Maybe it should be? The "dieselgate" talk[1] at 32c3 suggests engineering has gotten very good[2] at "teaching machines to the test".

[1] https://media.ccc.de/v/32c3-7331-the_exhaust_emissions_scand... (good text summary: https://lwn.net/Articles/670488/ )

[2] https://static.lwn.net/images/2016/vw-curves.png

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#42
post #29

Earlier quoted context omitted.

For #2, my immediate read was that the planes had been shot down. If the context were to suggest that the enemy had somehow hijacked the planes, then of course the word land would mean the same in both sentences. I have never used or heard 'land a plane' in this context, but the sentence didn't immediately strike me as unnatural, incorrect or unclear.

> I have never used or heard 'land a plane' in this context, but the sentence didn't immediately strike me as unnatural, incorrect or unclear. It struck me as pretty awkward and very ambiguous. It probably means 'obtained' but 'captured' would be a far better word in that case. The suggestions that it means 'hit/shot' don't work because in that case it's not the aircraft that is landed but the shot, which is landed o…

Time matters too. Current tech would hint to us that the planes had been shot down.

But in the future that sentence might mean hacking and theft of the actual planes, an actual landing.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#43
post #16

There was an article[1] posted to HN recently about these benchmarks, and it was pretty skeptical. Regarding SuperGLUE specifically, it asked: "Indeed, Bowman and his collaborators recently introduced a test called SuperGLUE that's specifically designed to be hard for BERT-based systems. So far, no neural network can beat human performance on it. But even if (or when) it happens, does it mean that machines can really…

This feels hollow. Can't this be said about any benchmark? It seems natural and proper that as one benchmark becomes saturated, we introduce harder benchmarks. I don't think anyone in the field thinks that once we match human performance on benchmark X, we're officially done. It just means it's time for more interesting benchmarks. Over time, if it starts to become difficult to design benchmarks that humans can outpe…

Yes. This is more generally known as Goodhart's Law[0]: when a metric is used as a goal, then people will game the metric in order to win, making the metric useless.

There is no fundamental way to overcome this problem, except by not using metrics as goals.

[0] https://en.wikipedia.org/wiki/Goodhart's_law

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#44

There was an article[1] posted to HN recently about these benchmarks, and it was pretty skeptical. Regarding SuperGLUE specifically, it asked: "Indeed, Bowman and his collaborators recently introduced a test called SuperGLUE that's specifically designed to be hard for BERT-based systems. So far, no neural network can beat human performance on it. But even if (or when) it happens, does it mean that machines can really…

Even when you will be able to have a 100% coherent and deep discussion with an AI over a niche technical domain, there will be people to pretend that the AI "fakes" it.

Systems like GPT-2, incredibly (I used to be a skeptic of a pure statistical approach) manage to extract meaning, keep a theme, and understand the intent behind a sentence. They are amazing.

When you have a system that displays all the characteristics of understanding something, it is irrelevant whether or not it "fakes" it. No one ever proved that humans are not "faking" intelligence either.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#45

Earlier quoted context omitted.

I've worked in aviation for 8 years and also didn't understand this use of "landed". I've heard "grounded" used like this: "The maintenance issues gounded the jet," but not "landed".

I think the sentence is referring to aircraft that have been forced to land by the enemy, in contrast to "grounded" aircraft that had not taken flight. I haven't worked in aviation so my understanding of terminology could be wrong, but either way it is definitely an unusual example.

Right, I think you understand the word as I do: 'verb' + ed. "The enemy landed the jet" as in they forced the jet to land either directly or indirectly. This would mean that the two sentences use "landed" the same way. But my understanding is SuperGLUE's offical answer is that these use "landed" differently with the rational that "landed" is idiomatic and just means to procure or bring about (e.g. "I landed the job") and it happens to be used with planes.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#46

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

The second one means "the enemy successfully got several of our aircrafts". Specifically, definition 3a or 3b for the verb form here: https://www.merriam-webster.com/dictionary/land So potentially the enemy captured the aircraft (3a) or destroyed them (3b).

If taking the "captured" interpretation, I think it could be reasonably inferred that they successfully landed the aircraft at an airfield afterwards (same meaning). This was my initial read of it and it does not seem strange to me on reflection.

I would like also to point out that even if we do interpret the second as meaning "destroyed", the first could then be interpreted as a combat aviator shooting down an opposing aircraft, bringing us back to the same meaning. Or perhaps both of my interpretations are correct and the meanings are different...

What this tells me is that the benchmark is not very useful.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#47
post #44

There was an article[1] posted to HN recently about these benchmarks, and it was pretty skeptical. Regarding SuperGLUE specifically, it asked: "Indeed, Bowman and his collaborators recently introduced a test called SuperGLUE that's specifically designed to be hard for BERT-based systems. So far, no neural network can beat human performance on it. But even if (or when) it happens, does it mean that machines can really…

Even when you will be able to have a 100% coherent and deep discussion with an AI over a niche technical domain, there will be people to pretend that the AI "fakes" it. Systems like GPT-2, incredibly (I used to be a skeptic of a pure statistical approach) manage to extract meaning, keep a theme, and understand the intent behind a sentence. They are amazing. When you have a system that displays all the characteristics…

[deleted]

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#48

Earlier quoted context omitted.

I think it's landed in the same sense as "landed a deal": got, or achieved, in this case achieving shooting them down. For me, my first read of the sentence would definitely be that it means shot down.

Ahh, just found an example where that's taken from https://glosbe.com/en/en/land . If you find on that page you'll see the exact sentence "the enemy landed several of our aircraft" (without the s after aircraft) which it says means "shoot down". I have still never heard landed used in that way, and again in other dictionaries I searched I couldn't find that definition either. Thus, this is a case where the "AI" may g…

I think if we really looked at it, it likely comes from fishing where "to land" a fish means to succeed in quite literally getting it onto land from the water. But we use it as "to successfully get" (something typically uncertain) in many other contexts.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#49

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

The second implies that the aircrafts were shot down; the first states that the aircraft landed safely. It looks like this reduces to the machine being able to figure out whether or not something is good or bad for the speaker.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#50

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

From the abstract of the associated paper: "performance on the benchmark has recently surpassed the level of non-expert humans, suggesting limited headroom for further research."

It occured to me that hn_throwaway_99's question, and the responses to it, is the sort of dialog in which one could find additional headroom for further research into natural language understanding. We can understand, for example, that while the two uses of 'landed' are different, they are not completely unrelated, and we can explain how they are related, for example by introducing a third construct, 'landed a fish', as a couple of replies have done.

Post reply on HN