Live data from Hacker News

ChatGPT outperforms crowd-workers for text-annotation tasks

arxiv.org

31–40 of 206 posts

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#31
post #12

Earlier quoted context omitted.

Also "the per-annotation cost of ChatGPT is less than $0.003 -- about twenty times cheaper than MTurk." It's interesting that the best available MTurk Master crowd-workers located in the US are paid about six cents per task.

I guess my surprise is that the machine is only 20x cheaper than the cheapest human available.

It’s worse because a US-based MTurker is far from the cheapest human available.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#32
post #20

My main take away here is that Turkers are terrible at some of these tasks. The "stance" task is, "Classify the tweet as having a positive stance towards Section 230, a negative stance, or a neutral stance.", and the Turkers accuracy was like 20%. Even in its best task, ChatGPT only got 75% accuracy.

20% means you can beat ChatGPT if you invert the answers!

There are 3 stances. But 33% is achievable by chance.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#33
post #21

Earlier quoted context omitted.

We have artificial intelligence that is general and above average human intelligence for the majority of tasks it can perform. Near expert level for some. NLP is a solved problem. Bespoke models are out the door. Large enough LLMs crush anything else for any NLP task. Honestly, this whole "they are not intelligent" argument is becoming ridiculous. might as well argue that a plane isn’t a real bird or a car isn’t a re…

> We have artificial intelligence that is general and above average human intelligence for the vast majority of tasks it can perform. Even when I give it the benefit of the doubt, this sentence makes no sense to me. Do we have a list of tasks a language model can perform? To the best of my knowledge, they can arguably perform any language task. > Large enough LLMs crush anything else for any NLP task. and evidently t…

You said NLP is unsolved because we don't have human level artificial intelligence. We absolutely do. at least by any evaluations we can carry out.

no-one wants to call a spade a spade yet but the sentiment is obvious in recent research. directly being called General purpose technologies from the jobs paper, general artificial intelligence from the creativity paper. That last one is particularly funny, they just switched the two words.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#34
post #9

Earlier quoted context omitted.

Idiomatic translation of text matching a human professional (e.g. free of errors for legal terms, interesting and natural for fiction) is unlikely to be achieved until we have AGI. So no.

I won't comment on the first bit as i've not personally tested in that are but GPT-4 can absolutely make short work on the second. I don't think people realize how good Bilingual LLMs are at translations. Yes you have idioms transfer between languages. Feel free to test it yourself.

What language pairs are you talking about? I don't think people realize just how much the difficulty level and the state of technology differ depending on that choice.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#36
post #5

Earlier quoted context omitted.

Obviously not. If we had a general solution to language intelligence we would have artificial intelligence at the level of at least human intelligence – which we do not . Rather, the right question to ask is which language intelligence tasks currently have acceptable performance and under which conditions (text domain, etc.). Clearly this is a much more difficult question and with a lot more nuance to it, even if it…

We have artificial intelligence that is general and above average human intelligence for the majority of tasks it can perform. Near expert level for some. NLP is a solved problem. Bespoke models are out the door. Large enough LLMs crush anything else for any NLP task. Honestly, this whole "they are not intelligent" argument is becoming ridiculous. might as well argue that a plane isn’t a real bird or a car isn’t a re…

> might as well argue that a plane isn’t a real bird or a car isn’t a real horse.

They aren’t though… They are far superior at specific things birds and horses are known for, but they can’t do everything that birds and horses can, so they aren’t even artificial birds and horses.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#37
post #17

Earlier quoted context omitted.

I have tested it :) I've asked it to translate English fictional text into Japanese, it falls over often. It's unnatural and often makes no sense at all. It doesn't compare to a typical professional translation (which are often not that idiomatic either), let alone a really good one. I'm sure it'll be doing that in five years, but not now. One interesting thing is that's it's nondeterministic, so sometimes 'For chris…

Are you giving it multiple paragraphs to translate at once so that it has enough context for a good translation? If so, would you mind sharing a sample input and output that you found unsatisfactory? In "Can GPT-4 translate literature?" (Mar 18, 2023) [ https://youtu.be/5KKDCp3OaMo?t=377 ], Tom Gally, a former professional translator and current professor at the University of Tokyo, said: > …the point is, to my eye a…

I don't think we disagree. The video says the translation will be "readable" but needs several days of an experienced editor passing over it. That's an amazing result, but again, it's not as good as a human yet. It's way faster and it'll make media accessible to tons of people.

Like he says, there's lots of ambiguity in Japanese that needs to be handled, gender not being specified until later, etc. and an editor would need to spend time going over it - but it saves months of traditional work. There are words and _concepts_ that are hard to translate, there are cultural issues, dialects, slang, registers. So yeah it'll make the media accessible, but it won't be as a good as a skilled translator.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#38

So is this 'AI trains AI better than people can'? And presumably the better-trained AI will also be better again at training. I think I've seen this movie.

That's exactly what's coming. AI will train AI. At some point AI is going to stop needing us to keep evolving. Not sure what that will be like.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#39
post #9

Earlier quoted context omitted.

Idiomatic translation of text matching a human professional (e.g. free of errors for legal terms, interesting and natural for fiction) is unlikely to be achieved until we have AGI. So no.

Which is to say that there are edge cases like legal texts or other fields where a high level of domain expertise is needed to interpret and translate text. Which most human translators would also not have. For almost everything else, it seems to produce pretty decent and usable translations, even when used against relatively obscure languages. I used it a some green landic article that was posted on hn yesterday (ab…

I can't vouch for the correctness obviously.

Therein lies the rub. There's a huge gap between what LLMs can currently do (spit back something in a target language that gives you the basic idea, however awkwardly phrased, of what was said in the source language). And what is actually needed for idiomatic, reasonably error-free translation.

By "reasonably error-free" I mean, say, requiring a human correction for less than 5 percent of all sentences. Current LLMs are nowhere near that level, even for resource-rich language pairs.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#40
post #34

Earlier quoted context omitted.

I won't comment on the first bit as i've not personally tested in that are but GPT-4 can absolutely make short work on the second. I don't think people realize how good Bilingual LLMs are at translations. Yes you have idioms transfer between languages. Feel free to test it yourself.

What language pairs are you talking about? I don't think people realize just how much the difficulty level and the state of technology differ depending on that choice.

English/Chinese is what i've tried on.

and you can see talk on English/Japanese here - https://youtu.be/5KKDCp3OaMo?t=377

Post reply on HN