Live data from Hacker News

ChatGPT outperforms crowd-workers for text-annotation tasks

arxiv.org

21–30 of 206 posts

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#21
post #5

Earlier quoted context omitted.

Obviously not. If we had a general solution to language intelligence we would have artificial intelligence at the level of at least human intelligence – which we do not . Rather, the right question to ask is which language intelligence tasks currently have acceptable performance and under which conditions (text domain, etc.). Clearly this is a much more difficult question and with a lot more nuance to it, even if it…

We have artificial intelligence that is general and above average human intelligence for the majority of tasks it can perform. Near expert level for some. NLP is a solved problem. Bespoke models are out the door. Large enough LLMs crush anything else for any NLP task. Honestly, this whole "they are not intelligent" argument is becoming ridiculous. might as well argue that a plane isn’t a real bird or a car isn’t a re…

> We have artificial intelligence that is general and above average human intelligence for the vast majority of tasks it can perform.

Even when I give it the benefit of the doubt, this sentence makes no sense to me. Do we have a list of tasks a language model can perform? To the best of my knowledge, they can arguably perform any language task.

> Large enough LLMs crush anything else for any NLP task. and evidently they beat top humans too.

Yes, they are certainly (rightfully) the go-to model for most tasks at this point if your concern is outright performance. Have I indicated otherwise? As for beating “top” humans, I am sure that can be investigated, but it is a fairly nuanced research question. It is inarguable that they are amazingly good though, especially relative to what we had just a few years ago.

> Honestly, this whole "they are not intelligent" argument is becoming ridiculously obtuse. > > might as well argue that a plane isn’t a real bird or a car isn’t a real horse.

Which is a claim and argument that I never made – hallucinating? How about you calm down a little and get back on the ground? You are talking to someone that has argued in favour of these kinds of models for about a decade. But that does not mean that I am willing to spout nonsense or lose track of what we know and what we do not yet know.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#22

So is this 'AI trains AI better than people can'? And presumably the better-trained AI will also be better again at training. I think I've seen this movie.

but its a super-spell checker in a way.. it does not understand the meaning of the patterns.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#23
post #17

Earlier quoted context omitted.

I won't comment on the first bit as i've not personally tested in that are but GPT-4 can absolutely make short work on the second. I don't think people realize how good Bilingual LLMs are at translations. Yes you have idioms transfer between languages. Feel free to test it yourself.

I have tested it :) I've asked it to translate English fictional text into Japanese, it falls over often. It's unnatural and often makes no sense at all. It doesn't compare to a typical professional translation (which are often not that idiomatic either), let alone a really good one. I'm sure it'll be doing that in five years, but not now. One interesting thing is that's it's nondeterministic, so sometimes 'For chris…

Mind sharing output ?

I mean i can if you want (chinese though) but enough people lie on the internet.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#24
post #17

Earlier quoted context omitted.

I have tested it :) I've asked it to translate English fictional text into Japanese, it falls over often. It's unnatural and often makes no sense at all. It doesn't compare to a typical professional translation (which are often not that idiomatic either), let alone a really good one. I'm sure it'll be doing that in five years, but not now. One interesting thing is that's it's nondeterministic, so sometimes 'For chris…

Are you giving it multiple paragraphs to translate at once so that it has enough context for a good translation? If so, would you mind sharing a sample input and output that you found unsatisfactory? In "Can GPT-4 translate literature?" (Mar 18, 2023) [ https://youtu.be/5KKDCp3OaMo?t=377 ], Tom Gally, a former professional translator and current professor at the University of Tokyo, said: > …the point is, to my eye a…

[deleted]

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#25
post #12
post #7

Before you ask, because I was curious, from the paper: "For MTurk, we aimed to select the best available crowd-workers, notably by filtering for workers who are classified as “MTurk Masters” by Amazon, who have an approval rate of over 90%, and who are located in the US."

Also "the per-annotation cost of ChatGPT is less than $0.003 -- about twenty times cheaper than MTurk." It's interesting that the best available MTurk Master crowd-workers located in the US are paid about six cents per task.

I guess my surprise is that the machine is only 20x cheaper than the cheapest human available.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#26
post #9
post #3

Is NLP a solved problem now?

Idiomatic translation of text matching a human professional (e.g. free of errors for legal terms, interesting and natural for fiction) is unlikely to be achieved until we have AGI. So no.

Which is to say that there are edge cases like legal texts or other fields where a high level of domain expertise is needed to interpret and translate text. Which most human translators would also not have.

For almost everything else, it seems to produce pretty decent and usable translations, even when used against relatively obscure languages.

I used it a some green landic article that was posted on hn yesterday (about Greenland having gotten rid of daylight saving time). I don't speak a word of that language but the resulting English translation looked like it matched the topic and generally read like correct and sensible English. I can't vouch for the correctness obviously. But I could not spot any weird errors or strange formulations that e.g. Google translate suffers from. That matches my earlier experience trying to get chat gpt to answer in some Dutch dialects, Frysian, Latin, and a few other more obscure outputs. It does all of that. Getting it to use pirate speak is actually quite funny.

The reason that I used Chat GPT for this is that Google translate does not understand greenlandic. Understandable because there are only a few tens of thousands of native speakers of that language and presumably there's not a very large amount of training material in that language.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#27
post #3

Is NLP a solved problem now?

I still think we're a long ways off. LLMs can't to my knowledge process a request into a lookup on say an actual database of facts at the moment or parse a request into API actions. So far it's shown it's really really good at continuing a conversation with more text but as far as I understand them there's not a usable comprehension of what's actually being asked and answered.

The point that would say to me the LLM actually has any "understanding" of what it's saying would be when it's able to reliably say "I don't know the answer to that" instead of making up things from scratch. You see that a lot if you ask Bing/Bard "Who is _____?" Most of them are kind of right but a lot of large details are just completely fabricated. A lot of the facts it gets wrong are things Google is already able to produce when queried like where was Person X born or where did they go to school so the fact these LLMs can't slot in actual available facts says to me they're not really going to be that useful with the kind of tasks we've been working on NLP for.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#28
post #20

My main take away here is that Turkers are terrible at some of these tasks. The "stance" task is, "Classify the tweet as having a positive stance towards Section 230, a negative stance, or a neutral stance.", and the Turkers accuracy was like 20%. Even in its best task, ChatGPT only got 75% accuracy.

20% means you can beat ChatGPT if you invert the answers!

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#29

So is this 'AI trains AI better than people can'? And presumably the better-trained AI will also be better again at training. I think I've seen this movie.

but its a super-spell checker in a way.. it does not understand the meaning of the patterns.

[dead]
Post reply on HN