Live data from Hacker News

Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

super.gluebenchmark.com

161–170 of 246 posts

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#161

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

Probably: The enemy grounded several of our aircrafts

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#162

This clearly demonstrates once again that Google is miles ahead of the competition in AI. I mean, they just have the best data. If you want to have an every day example of Google's AI skills: Switch you phone's keyboard to GBoard, especially all iOS users, and you will face a night and day difference to any other keyboard esepcially the stock one. When using multiple languages at the same time the leap to other keybo…

Have you tried the iOS 13 keyboard's built in swipe feature?

Are you aware of Swiftkey?

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#163
post #117
post #96

Earlier quoted context omitted.

My mother got a perfect 800 score on the GRE English test many years ago when she wanted to go back to graduate school after her children were grown up enough (highschool/college age). She told me that the way she got her perfect score was by realizing when the questions were wrong and thinking of what answer the test creators believed to be correct. She had to outguess the test creators and answer the questions wron…

I've had the 'pleasure' of taking some 'Microsoft certifications' at various companies I worked at in the past and this sounds extremely familiar. "I probably won't ever do it like that and/or there's a syntax error in all four of the answers... but this is the answer you want to hear. It's wrong, mind you, but it's what you want to hear."

Yep! You have to do away with conventional logic and ask yourself "What insanity would Microsoft recommend I do?"

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#164

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

I think this is really interesting, because "the enemy landed several of our aircraft(s)" is the sort of sentence I'd have hauled a student up for using as a teacher, because 1) it's a none standard, arguably incorrect usage they've used either because they're a none native speaker or because they're trying to be clever and failing, and 2) because the plural of aircraft is aircraft. Nevertheless the author of this se…

If you teach others English, please learn the difference between "none" and "non". You mean "non-standard" in all your examples here (if British) or perhaps "nonstandard" (if American).

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#165

As someone working in the field, I congratulate the excellent accomplishment but agree with the authors that we shouldn't get too excited yet (their quote below after the four reasons). Here are some reasons: 1) Most likely, the model is still susceptible to adversarial triggers as demonstrated on other systems here: http://www.ericswallace.com/triggers 2) T5 was trained with ~750GB of texts or ~150 billion words, wh…

> 2) T5 was trained with ~750GB of texts or ~150 billion words, which is > 100 times the number of words native English speakers acquire by the age of 20. ...but, humans evolved the ability to use language over hundreds of generations... So... Maybe that's not such a bad thing?

Indeed this is important to realize: Training such a generic model from scratch does not only reiterate learning, but the entire evolutionary process that led to the emergence of neural circuits actually capable of such learning. That perspective makes many of the current achievements -- error-prone as they might be -- even more impressive!

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#166

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

The plural of aircraft is aircraft. Not aircrafts. https://www.grammar-monster.com/plurals/plural_of_aircraft.h... Maybe the _examples_ for a language test should be grammatically correct?

It depends on what your goal is. But in most cases, I'd say no. If the goal has anything to do with understanding real language written by real humans, it's better for the system to be able to handle texts with errors.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#167
post #96

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

My mother got a perfect 800 score on the GRE English test many years ago when she wanted to go back to graduate school after her children were grown up enough (highschool/college age). She told me that the way she got her perfect score was by realizing when the questions were wrong and thinking of what answer the test creators believed to be correct. She had to outguess the test creators and answer the questions wron…

That's how I got through the SAT…

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#168

Earlier quoted context omitted.

Working in aviation probably puts you in a mindset that makes it harder to parse. It's not being used in a way that is related to flight or aircraft. It's like if people were discussing where to have a conference, and one of them proposed a hotel. Then another person suggested a resort. Then a third person floated a cruise ship. Cruise ships do float, but it has nothing to do with anything. They are floating the idea…

Do you normally "float" a cruise ship though? A more apt analogy might be "dock". Maybe a news report says that a vacation company has broken some regulation so the government docked a cruise ship, meaning they took away a cruise ship like you would dock someone points. It's ambiguous at best.

You could float the idea of it, and you might also think that to float a ship means the process by which it is landed in the water when coming out of a dock?

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#169
post #130

Earlier quoted context omitted.

I think it's archaic; in the past a fishing reference would have been more common and widely understood.

I guess it annoys me because I suspect that if this is the sort of borderline incoherence one must wade through I would probably score below average.

Or just average. There's contextual dependencies in most speech, and (as displayed in this subthread) not every speaker of a language has the same context. It's a fallacy to think that if you lack context for one of the examples, you will automatically score less than average -- other people may miss context for things obvious to you.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#170
post #96

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

My mother got a perfect 800 score on the GRE English test many years ago when she wanted to go back to graduate school after her children were grown up enough (highschool/college age). She told me that the way she got her perfect score was by realizing when the questions were wrong and thinking of what answer the test creators believed to be correct. She had to outguess the test creators and answer the questions wron…

I have achieved similar results by similar means in both English and certain other subjects wherein one would assume a “true academic” would “know better” (picking out Sin[x]=2 as being “evidence of error in prior working” when x could merely be Complex, or marking “f[f[n]]=-n as “unsolvable” when it’s just requires a bit of lateral thinking). This always depresses me, like when (as a Brit) I hear Americans say “I could care less” as an indicator of disregard, when actually that indicates they are somewhere above the point of minimal regard.
Post reply on HN