Live data from Hacker News

Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

super.gluebenchmark.com

171–180 of 246 posts

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#171

Earlier quoted context omitted.

I've never seen "landed" used as in the second sentence, but I was definitely able to understand from context that it was not being used to mean the same thing as in the first sentence.

You've never landed a fish?

Is "landing a fish" the same thing as "watering a plant"?

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#172
This surprised me a bit, on the creation of the corpus they use for training:

"We removed any page that contained any word on the “List of Dirty, Naughty, Obscene or Otherwise Bad Words”."

I don't understand this decision. This list contains words that can be used in a perfectly objective sense, like "anus", "bastard", "erotic", "eunuch", "fecal", etc.

I can understand that they want to avoid websites full of expletives and with no useful content, but outright excluding any website with even one occurrence of such words sounds too radical. If we ask this model a text comprehension question about a legitimized bastard that inherited the throne, or about fecal transplants, I suppose it would easily fail. Strange way of limiting such a powerful model.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#173

Earlier quoted context omitted.

The second one means "the enemy successfully got several of our aircrafts". Specifically, definition 3a or 3b for the verb form here: https://www.merriam-webster.com/dictionary/land So potentially the enemy captured the aircraft (3a) or destroyed them (3b).

Would a native English speaker use the word "landed" in this way? In the context of aircraft? "Landed" is badly ambiguous here and several distinct meanings are plausible. Captured is the most natural word given your interpretation. Honestly that sentence -- the use of landed and that awful plural -- approaches engrish. Is that deliberate or is the use of English here just badly flawed? I can't see any other possibil…

I don't think anyone would use that particular construction, unless it's some weird dialect of pilot-speak or argot among anti-aircraft folk that I'm not aware of. It's just really awkward and unnatural. Possibly correct, but not the way that anybody actually talks.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#174
post #54

Earlier quoted context omitted.

Limited headroom? Seems like they're assuming greater-than-human language ability is just impossible and will never be surpassed.

I'd argue that greater-tham-human language ability is by definition useless. Language is specifically a human communication tool, there's no value in surpassing the language skill that humans have, if indeed such a thing is even meaningful (what does it mean to be better than the best* French person at French?) * By whatever language-related metric

I disagree, greater-than-human-average is not useless. There's a lot of room for misinterpretation in human language. We compensate for that by non-verbal communication (posture, expression) or by asking for clarification. On top of that, most places have local expressions or idioms that are not necessarily globally recognized.

So there's two ways in which a language automaton must be better than human: it cannot rely on non-verbal hints nor can it easily ask for clarification, and it must be able to interpret many different dialects and idioms correctly -- many more than an average human would need to.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#175
post #96

Earlier quoted context omitted.

My mother got a perfect 800 score on the GRE English test many years ago when she wanted to go back to graduate school after her children were grown up enough (highschool/college age). She told me that the way she got her perfect score was by realizing when the questions were wrong and thinking of what answer the test creators believed to be correct. She had to outguess the test creators and answer the questions wron…

This does seem like the meta solution to most tests, particularly standardised tests :)

Paraphrasing Simonyi: “Any test you can pass, I can pass meta”.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#176

Earlier quoted context omitted.

I think this is really interesting, because "the enemy landed several of our aircraft(s)" is the sort of sentence I'd have hauled a student up for using as a teacher, because 1) it's a none standard, arguably incorrect usage they've used either because they're a none native speaker or because they're trying to be clever and failing, and 2) because the plural of aircraft is aircraft. Nevertheless the author of this se…

If you teach others English, please learn the difference between "none" and "non". You mean "non-standard" in all your examples here (if British) or perhaps "nonstandard" (if American).

(And non-native)

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#177
post #98

I think classifying this as human level is misleading. Look at the sub-scores on the page. One score that looks very different from humans is AX-b. The SuperGlue paper provides more context about AX-b https://arxiv.org/pdf/1905.00537.pdf AX-b "is the broad-coverage diagnostic task, scored using Matthews’ correlation (MCC). " This is how the paper describes this test " Analyzing Linguistic and World Knowledge in Model…

Hi, one of the paper's authors here. We didn't submit our model's predictions for the AX-b task yet, we just copied over the predictions from the example submission. We will submit predictions for AX-b in the next few days.

RcouF1uZ4gsC makes a compelling case for the results on this test to potentially be a significant caveat to the results, and also to the claims of achieving a near-human level of performance. If so, then why would you make such claims before you have these results? Or at least mention this caveat at the points where you are making the claim, such as in the abstract.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#178
post #163
post #117

Earlier quoted context omitted.

I've had the 'pleasure' of taking some 'Microsoft certifications' at various companies I worked at in the past and this sounds extremely familiar. "I probably won't ever do it like that and/or there's a syntax error in all four of the answers... but this is the answer you want to hear. It's wrong, mind you, but it's what you want to hear."

Yep! You have to do away with conventional logic and ask yourself "What insanity would Microsoft recommend I do?"

It's not always insanity, sometimes just sub-optimal / way over-engineered in my opinion.

They're getting better at it though. More recently I've done their devops certification and it looks like they're recommending somewhat more sane practices now...

There were still questions where even after three or four tries at certification / reading up on whatever Microsoft thinks is 'good' we didn't find 'the correct answer' according to Microsoft though... ¯\_(ツ)_/¯

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#179
post #178
post #163

Earlier quoted context omitted.

Yep! You have to do away with conventional logic and ask yourself "What insanity would Microsoft recommend I do?"

It's not always insanity, sometimes just sub-optimal / way over-engineered in my opinion. They're getting better at it though. More recently I've done their devops certification and it looks like they're recommending somewhat more sane practices now... There were still questions where even after three or four tries at certification / reading up on whatever Microsoft thinks is 'good' we didn't find 'the correct answer…

Yeah, that's true. It's still a good idea to get an idea of what a desired answer would be, which is why those answer dumps are so popular.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#180

This clearly demonstrates once again that Google is miles ahead of the competition in AI. I mean, they just have the best data. If you want to have an every day example of Google's AI skills: Switch you phone's keyboard to GBoard, especially all iOS users, and you will face a night and day difference to any other keyboard esepcially the stock one. When using multiple languages at the same time the leap to other keybo…

I have the opposite experience. Yes, some of the suggestions from GBoard are useful, but I feel there's an equal number of times where I've typed a complete word, only to hit space and have the word auto-corrected to what GBoard was expecting. As a typing aid, it's almost unusable because of that.
Post reply on HN