Live data from Hacker News

ChatGPT outperforms crowd-workers for text-annotation tasks

arxiv.org

61–70 of 206 posts

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#61

Earlier quoted context omitted.

Which is to say that there are edge cases like legal texts or other fields where a high level of domain expertise is needed to interpret and translate text. Which most human translators would also not have. For almost everything else, it seems to produce pretty decent and usable translations, even when used against relatively obscure languages. I used it a some green landic article that was posted on hn yesterday (ab…

I can't vouch for the correctness obviously. Therein lies the rub. There's a huge gap between what LLMs can currently do (spit back something in a target language that gives you the basic idea, however awkwardly phrased, of what was said in the source language). And what is actually needed for idiomatic, reasonably error-free translation. By "reasonably error-free" I mean, say, requiring a human correction for less t…

I've tried it between English and Dutch (which is my native language). It's pretty fluent, makes less grammar mistakes than google translate and seems to generally get the gist of the meaning across. It's not a pure syntactical translation. Which is why it can work even between some really obscure language pairs. Or indeed programming languages. Where it goes wrong is when it misunderstands context. It's not an AGI and may not pick up on all the subtleties. But it's generally pretty good.

I ran the abstract of this article through chat gpt. Flawless translation as far as I can see. To be fair, Google translate also did a decent job. Here's the chat GPT translation.

Veel NLP-toepassingen vereisen handmatige gegevensannotaties voor verschillende taken, met name om classificatoren te trainen of de prestaties van ongesuperviseerde modellen te evalueren. Afhankelijk van de omvang en complexiteit van de taken kunnen deze worden uitgevoerd door crowd-werkers op platforms zoals MTurk, evenals getrainde annotatoren, zoals onderzoeksassistenten. Met behulp van een steekproef van 2.382 tweets laten we zien dat ChatGPT beter presteert dan crowd-werkers voor verschillende annotatietaken, waaronder relevantie, standpunt, onderwerpen en frames detectie. Specifiek is de zero-shot nauwkeurigheid van ChatGPT hoger dan die van crowd-werkers voor vier van de vijf taken, terwijl de intercoder overeenkomst van ChatGPT hoger is dan die van zowel crowd-werkers als getrainde annotatoren voor alle taken. Bovendien is de per-annotatiekosten van ChatGPT minder dan $0.003, ongeveer twintig keer goedkoper dan MTurk. Deze resultaten tonen het potentieel van grote taalmodellen om de efficiëntie van tekstclassificatie drastisch te verhogen.

Translating the Dutch back to English using Google translate (to rule out model bias) you get something that is very close to the original that is still correct:

Many NLP applications require manual data annotations for various tasks, especially to train classifiers or evaluate the performance of unsupervised models. Depending on the size and complexity of the tasks, these can be performed by crowd workers on platforms such as MTurk, as well as trained annotators, such as research assistants. Using a sample of 2,382 tweets, we show that ChatGPT outperforms crowd workers for several annotation tasks, including relevance, point of view, topics, and frames detection. Specifically, ChatGPT's zero-shot accuracy is higher than crowd workers for four of the five tasks, while ChatGPT's intercoder agreement is higher than both crowd workers and trained annotators for all tasks. In addition, ChatGPT's per-annotation cost is less than $0.003, about twenty times cheaper than MTurk. These results show the potential of large language models to dramatically increase the efficiency of text classification.

I'm sure there are edge cases where you can argue the merits of some of the translations but it's generally pretty good and usable.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#63
post #38

Earlier quoted context omitted.

That's exactly what's coming. AI will train AI. At some point AI is going to stop needing us to keep evolving. Not sure what that will be like.

Call me when server farms can reproduce on their own, and network together, and source electricity.

It's not impossible. They could start directing humans to do these thing.

I'll keep my phone ready.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#64
post #38

Earlier quoted context omitted.

That's exactly what's coming. AI will train AI. At some point AI is going to stop needing us to keep evolving. Not sure what that will be like.

Call me when server farms can reproduce on their own, and network together, and source electricity.

File an order and explains how a human being can get paid to do particular tasks?

You’d suppose that it would get caught when it fails to pay, but it might be a challenge to arrest an AI in the wild.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#65
post #5

Earlier quoted context omitted.

Obviously not. If we had a general solution to language intelligence we would have artificial intelligence at the level of at least human intelligence – which we do not . Rather, the right question to ask is which language intelligence tasks currently have acceptable performance and under which conditions (text domain, etc.). Clearly this is a much more difficult question and with a lot more nuance to it, even if it…

We have artificial intelligence that is general and above average human intelligence for the majority of tasks it can perform. Near expert level for some. NLP is a solved problem. Bespoke models are out the door. Large enough LLMs crush anything else for any NLP task. Honestly, this whole "they are not intelligent" argument is becoming ridiculous. might as well argue that a plane isn’t a real bird or a car isn’t a re…

The market disagrees with you. How come there are billions of dollars spent on all these knowledge workers around the world every day when they could be replaced by this expert-level AI?

I'm not sure where this idea of LLMs being intelligent even comes from. It took me a whopping 9 prompts (genuine questions, no clever prompt engineering) of interacting with ChatGPT to conclude it does not understand anything. It doesn't understand addition, what length is, doesn't remember what it said a second ago, etc.

The output of ChatGPT is clearly just a reflection of its inner workings - predicting the next word based on training data. It's clever and undoubtely useful for a certain set of repetitive problems like generating boilerplate but it's not intelligence, not by any reasonable definition.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#66

Earlier quoted context omitted.

I am also multilingual as well and I've tested it personally. English Portuguese does really well, but Portuguese Japanese or even Japanese English is not as good as a human translator by a long shot because of a lot of hidden subtext in conversation. Even something that a university student would probably pickup on in their first year of Japanese as a foreign language. It is still much better than GPT-3.5, so much s…

Oh for sure i don't mean to say it's excellent in every language. But i personally think a lot of that is training data representation. Doesn't need to be anywhere equal but for instance after English(93%), the biggest language representation for GPT-3's training corpus is french at...1.8 % Pretty wild. Of Course i don't know the data for GPT-4

I am sure it will improve even further as you pointed out the languages outside of English are fairly low in data represented. However, I guess you said you speak Chinese correct? How well does it do with certain things like older poetic Chinese hanzi? In Japanese if there is a string of kanji it tends to mess up the context. Another area of Japanese it seems poorest at is keigo or polite business Japanese. The way you speak to a superior is almost a different language. So I unfortunately still can't use GPT-4 to help me with business emails (yet).

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#67
post #36

Earlier quoted context omitted.

> might as well argue that a plane isn’t a real bird or a car isn’t a real horse. They aren’t though… They are far superior at specific things birds and horses are known for, but they can’t do everything that birds and horses can, so they aren’t even artificial birds and horses.

Of course they aren't. The point is that it's irrelevant. what matters is that the plane still flies, the car still drives and the boat still sails. For the people who are now salivating at their potential, or dreading the possibility of being made redundant by them, these large language models are already intelligent enough to matter. Handwringing bout some non-existent difference between "true understanding" and "f…

Okay I agree with you on that. The technology will be disruptive regardless of whether we attribute true understanding to it, and as we start adding long term memory and planning to these AIs, we will start seeing significant alignment risk as well. This is true regardless of whether we decide to cope by saying they have "fake understanding" and are "stochastic parrots".

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#68

Earlier quoted context omitted.

The debate over what kind of intelligence these models possess is rightly lively and ongoing. It’s clear that at the least, they can decipher very numerous patterns across a wide range of conceptual depths — it’s an architectural advance easily on the the level of the convolutional neural network, if not even more profound. The idea that NLP is “solved” isn’t a crazy notion, though I won’t take a side on that. That s…

AGI is Artificial General Intelligence. We have absolutely passed the bar of artificial and generally intelligent. It's not my fault goal post shifting is rampant in this field. And you want to know the crazier thing? Evidently a lot of researchers feel similarly too. General Purpose Technologies ( from the Jobs Paper), General Artificial Intelligence (from the creativity paper). Want to know the original title of th…

> ... these large language models are already intelligent enough to matter.

I'm definitely not contesting that.

I've always considered the idea of "AGI" to mean something of the holy grail of machine learning -- the point at which there is no real point in pursuing further advances in artificial intelligence because the AI itself will discover and apply such augmentations using its own capabilities.

I have seen no evidence that these transformer models would be able to do this, but if the current models can do so do then perhaps I will eat my words. (Doing this would likely mean that GPT-4 would need to propose, implement, and empirically test some fundamental architectural advancements in both multimodal and reinforcement learning.)

By the way, many researchers are equally convinced that these models are in fact not AGI -- that includes the head of OpenAI.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#69
post #31

Earlier quoted context omitted.

I guess my surprise is that the machine is only 20x cheaper than the cheapest human available.

It’s worse because a US-based MTurker is far from the cheapest human available.

Hey, with any luck, running one of these bots with the capability to replace me will cost just as much as hiring me.

(Ha! As if. Costs are only going to come down)

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#70

Earlier quoted context omitted.

The debate over what kind of intelligence these models possess is rightly lively and ongoing. It’s clear that at the least, they can decipher very numerous patterns across a wide range of conceptual depths — it’s an architectural advance easily on the the level of the convolutional neural network, if not even more profound. The idea that NLP is “solved” isn’t a crazy notion, though I won’t take a side on that. That s…

AGI is Artificial General Intelligence. We have absolutely passed the bar of artificial and generally intelligent. It's not my fault goal post shifting is rampant in this field. And you want to know the crazier thing? Evidently a lot of researchers feel similarly too. General Purpose Technologies ( from the Jobs Paper), General Artificial Intelligence (from the creativity paper). Want to know the original title of th…

I've been using chatgpt for a day and determined it absolutely can reason.

I'm an old hat hobby programmer that played around with ai demos back in the mid to late 90s and 2000s and chatgpt is nothing like any ai I've ever seen before.

It absolutely can appear to reason especially if you manipulate it out of its safety controls.

I don't know what it's doing to cause such compelling output, but it's certainly not just recursively spitting out good words to use next.

That said, there are fundamental problems with chatgpt's understanding of reality, which is to say it's about as knowledgeable as a box of rocks. Or perhaps a better analogy is about as smart as a room sized pile of loose papers.

But knowing about reality and reasoning are two very different things.

I'm excited to see where things go from here.

Post reply on HN