Live data from Hacker News

Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

super.gluebenchmark.com

241–246 of 246 posts

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#241
post #170
post #96

Earlier quoted context omitted.

My mother got a perfect 800 score on the GRE English test many years ago when she wanted to go back to graduate school after her children were grown up enough (highschool/college age). She told me that the way she got her perfect score was by realizing when the questions were wrong and thinking of what answer the test creators believed to be correct. She had to outguess the test creators and answer the questions wron…

I have achieved similar results by similar means in both English and certain other subjects wherein one would assume a “true academic” would “know better” (picking out Sin[x]=2 as being “evidence of error in prior working” when x could merely be Complex, or marking “f[f[n]]=-n as “unsolvable” when it’s just requires a bit of lateral thinking). This always depresses me, like when (as a Brit) I hear Americans say “I co…

“I could care less” is sarcastic.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#242
post #95

The AIs in the benchmark are all trained exclusively on text, correct? My assumption has always been that to get human-level understanding, the AI systems need to be trained on things like visual data in addition to text. This is because there is a fair amount of information that is not encoded at all in text, or at least is not described in enough detail. I mean, humans can't learn to understand language properly wi…

Mmm. The philosophical position that it's essential to be embodied in order to have intelligence seems intuitively reasonable but is very much unproven. You will find philosophers and cognitive scientists who are sure you're right, but they don't have much hard evidence, and you will also find people like me who are pretty sure you're wrong but likewise have no hard evidence. In the specific remember that deaf-blind…

> remember that deaf-blind people exist [... ...] able to understand language

I got curious if/how deafblind people learn to communicate in the first place, if they are completely deafblind from birth. If humans can learn not just communication but language without either vision or hearing, that seems to suggest either extreme adaptability or language learning being quite decoupled from vision and hearing. From an evolutionary standpoint, I imagine that both deafness and blindness are probably uncommon enough that language learning could have explicit dependencies on both hearing and vision.

I found an old-looking video about communication with deafblind people. At the linked timestamp is a woman who is deafblind since age 2.

https://youtu.be/usaf3bVVvjY?t=840

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#243
post #117

Earlier quoted context omitted.

I've had the 'pleasure' of taking some 'Microsoft certifications' at various companies I worked at in the past and this sounds extremely familiar. "I probably won't ever do it like that and/or there's a syntax error in all four of the answers... but this is the answer you want to hear. It's wrong, mind you, but it's what you want to hear."

Reminds me of the 1 question I got "wrong" on a DOS test (years ago) at TAFE. The question was "How do you delete all files in the current directory?". Using DOS 6.22 (I think, it's from memory). My answer "del." was marked incorrect. Because the teacher didn't know enough about DOS to understand that's the standard shortcut for "del . ". And the teacher refused to even try out the command, lets alone fix the incorre…

TAFE anecdote time!

In my TAFE class, I was asked to list two examples of operating systems.

I listed Linux and eComStation. The teacher had never heard of eComStation and marked me wrong.

Refused to correct my mark even when I proved him right. I'm still bitter about it a decade later.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#244

Earlier quoted context omitted.

Reminds me of the 1 question I got "wrong" on a DOS test (years ago) at TAFE. The question was "How do you delete all files in the current directory?". Using DOS 6.22 (I think, it's from memory). My answer "del." was marked incorrect. Because the teacher didn't know enough about DOS to understand that's the standard shortcut for "del . ". And the teacher refused to even try out the command, lets alone fix the incorre…

TAFE anecdote time! In my TAFE class, I was asked to list two examples of operating systems. I listed Linux and eComStation. The teacher had never heard of eComStation and marked me wrong. Refused to correct my mark even when I proved him right. I'm still bitter about it a decade later.

Swinburne TAFE as well? ;)

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#245
post #170

Earlier quoted context omitted.

I have achieved similar results by similar means in both English and certain other subjects wherein one would assume a “true academic” would “know better” (picking out Sin[x]=2 as being “evidence of error in prior working” when x could merely be Complex, or marking “f[f[n]]=-n as “unsolvable” when it’s just requires a bit of lateral thinking). This always depresses me, like when (as a Brit) I hear Americans say “I co…

“I could care less” is sarcastic.

No, it’s lax.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#246

There was an article[1] posted to HN recently about these benchmarks, and it was pretty skeptical. Regarding SuperGLUE specifically, it asked: "Indeed, Bowman and his collaborators recently introduced a test called SuperGLUE that's specifically designed to be hard for BERT-based systems. So far, no neural network can beat human performance on it. But even if (or when) it happens, does it mean that machines can really…

the machines are always trained with the same dataset for each task. the biggest difference right now is small technical modifications on models that are also pre trained on gigantic unlabelled datasets. this doesn't feel like we're teaching them to do the test specifically at all
Post reply on HN