Live data from Hacker News

Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

super.gluebenchmark.com

181–190 of 246 posts

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#181

This clearly demonstrates once again that Google is miles ahead of the competition in AI. I mean, they just have the best data. If you want to have an every day example of Google's AI skills: Switch you phone's keyboard to GBoard, especially all iOS users, and you will face a night and day difference to any other keyboard esepcially the stock one. When using multiple languages at the same time the leap to other keybo…

Have you tried the iOS 13 keyboard's built in swipe feature? Are you aware of Swiftkey?

I think GP is talking about predictive text, rather than keyboard ergonomics.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#182

One thing to always point out in these cases is that the human baseline isn't "how well people do at this task," like it's often hyped to be. It's "how well does a person quickly and repetitively doing this do, on average." The 'quickly and repetitively' part is important because we all make more boneheaded errors in this scenario. The 'on average' part is important because the errors the algo makes aren't just fewer…

In the context of GPT2 someone coined the expression "Humans Who Are Not Concentrating Are Not General Intelligences"

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#183
post #95

The AIs in the benchmark are all trained exclusively on text, correct? My assumption has always been that to get human-level understanding, the AI systems need to be trained on things like visual data in addition to text. This is because there is a fair amount of information that is not encoded at all in text, or at least is not described in enough detail. I mean, humans can't learn to understand language properly wi…

Mmm. The philosophical position that it's essential to be embodied in order to have intelligence seems intuitively reasonable but is very much unproven. You will find philosophers and cognitive scientists who are sure you're right, but they don't have much hard evidence, and you will also find people like me who are pretty sure you're wrong but likewise have no hard evidence.

In the specific remember that deaf-blind people exist, so if you're sure that you "need something visual or auditory" then those people are not, according to your beliefs, able to understand language. I think they'll disagree with you quite strongly.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#184

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

I believe it’s “land” in the sense of “land a fish” (or a prize in general) which is a less common but legitimate usage.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#185

Earlier quoted context omitted.

I've worked in aviation for 8 years and also didn't understand this use of "landed". I've heard "grounded" used like this: "The maintenance issues gounded the jet," but not "landed".

I think the sentence is referring to aircraft that have been forced to land by the enemy, in contrast to "grounded" aircraft that had not taken flight. I haven't worked in aviation so my understanding of terminology could be wrong, but either way it is definitely an unusual example.

A fishing boat can land a big catch - and a sales executive might have landed a big deal, perhaps after reeling them in or having them on the hook.

So this would be particularly apt wording if the enemy had thrown a net over the plane as it sank in the ocean.

But I prefer to think the enemy gifted british country estates to the planes.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#186

I didn't know anything about SuperGLUE before (turns out it's a benchmark for language understanding tasks), so I clicked around their site where they show different examples of the tasks. One "word in context" task is to look at 2 different sentences that have a common word and decide if that word means the same thing in both sentences or different things (more details here: https://pilehvar.github.io/wic/ ) One of…

The example directly below that: "Justify the margins" and "The end justifies the means" is the one I find dubious. Obviously the former could mean to format a document, but those exact words in that structure could be a demand for someone to justify a financial margin for example. It is both true and false depending on the context.

I'm guessing this is intentional. To a human, although this could be somebody being asked to justify their financial margins that's not a very likely answer. The human can easily see that, while it's possible they're the same meaning, given the lack of any other context the answer is that they're not.

The enemy could have landed several of our aircraft on one of their runways. Agassi may have beaten Becker over the head with his tennis racket. I suspect part of the test is that there can be other meanings that do technically work.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#187
post #95

The AIs in the benchmark are all trained exclusively on text, correct? My assumption has always been that to get human-level understanding, the AI systems need to be trained on things like visual data in addition to text. This is because there is a fair amount of information that is not encoded at all in text, or at least is not described in enough detail. I mean, humans can't learn to understand language properly wi…

I think maybe CLEVR[0] dataset is what you are talking about? Keep in mind that a most of the current ML systems have diverged from biology. A majority of the recent breakthroughs come from mathematics, the rational is that just because human brain does it in a certain way does not necessarily mean it is the only way to do it. [0] https://cs.stanford.edu/people/jcjohns/clevr/

In no way do I think that AGI needs to mimic animal/human intelligence.

I was just trying to explain why text input alone isn't going to be adequate and that was an example.

Thanks for the link, that is one example of the type of thing I was talking about I think.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#188

Earlier quoted context omitted.

I think this is really interesting, because "the enemy landed several of our aircraft(s)" is the sort of sentence I'd have hauled a student up for using as a teacher, because 1) it's a none standard, arguably incorrect usage they've used either because they're a none native speaker or because they're trying to be clever and failing, and 2) because the plural of aircraft is aircraft. Nevertheless the author of this se…

If you teach others English, please learn the difference between "none" and "non". You mean "non-standard" in all your examples here (if British) or perhaps "nonstandard" (if American).

Sssh, pumpkin. We live in a world of autocorract.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#189
post #117
post #96

Earlier quoted context omitted.

My mother got a perfect 800 score on the GRE English test many years ago when she wanted to go back to graduate school after her children were grown up enough (highschool/college age). She told me that the way she got her perfect score was by realizing when the questions were wrong and thinking of what answer the test creators believed to be correct. She had to outguess the test creators and answer the questions wron…

I've had the 'pleasure' of taking some 'Microsoft certifications' at various companies I worked at in the past and this sounds extremely familiar. "I probably won't ever do it like that and/or there's a syntax error in all four of the answers... but this is the answer you want to hear. It's wrong, mind you, but it's what you want to hear."

Sometimes the questions are also just broken. I.e. asking you to select the things that do not apply, but the answer would have been to opposite.

Re: Google T5 scores 88.9 on SuperGLUE Benchmark, approaching human baseline

#190

One thing to always point out in these cases is that the human baseline isn't "how well people do at this task," like it's often hyped to be. It's "how well does a person quickly and repetitively doing this do, on average." The 'quickly and repetitively' part is important because we all make more boneheaded errors in this scenario. The 'on average' part is important because the errors the algo makes aren't just fewer…

In the context of GPT2 someone coined the expression "Humans Who Are Not Concentrating Are Not General Intelligences"

I think it was this blogger: https://www.google.com/amp/s/srconstantin.wordpress.com/2019...
Post reply on HN