Live data from Hacker News

ChatGPT outperforms crowd-workers for text-annotation tasks

arxiv.org

181–190 of 206 posts

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#181
post #179

Earlier quoted context omitted.

I have long suspected that Turkers dishonestly perform the tasks. At 0.06 cents per task, you're really incentivizing "finish the task as quickly as possible"; and "press the left button" is a lot faster than "read the tweet, think about it, and classify".

> At 0.06 cents per task Do you work at Verizon?

I got that figure from https://news.ycombinator.com/item?id=35335558

[append] Oh, it's the xkcd Verizon agent reference. Got it. My bad.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#182

Earlier quoted context omitted.

thanks for your answer. thats a reasonable point - but would we be at a tipping point by GPT5/6 (chatgpt is gpt 3.5) where human alignment is not needed? in fact, my question is reinforced by the GPT-4 technical report which explicitly mentioned that RLHF did NOT make a change to performance (and was only used for safety purposes)

GPT6 or whatever will always require alignment, as the base model just blindly predicts next token, instead of being a helpful, chat style assistant. Right now the best way to align it is with RLHF. The specific technique might change, but in the end there will always be at some level some human input that tells it how it should behave. Newer techniques might further leverage LLMs and require fewer human input. Could…

you don't need humans in the loop for alignment. rlaif is a thing and is used for the anthropic models (claude)

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#183
"ChatGPT’s intercoder agreement exceeds that of both crowd-workers and trained annotators for all tasks." -- wait, what? However good ChatGPT is at approximating trained annotators for these Twitter tasks, it's an algorithm, so the level of simulated "inter-annotator agreement" is in the authors' control (in the case of GPT, via the temperature parameter, for which they try just two values, 0.2 and 1.)

And why does this paper not make any effort to describe the wide range of annotation tasks for which this kind of simulated annotation is not a good idea -- for example, where you care about the subjective opinions of specific subgroups of people at specific times. And even for the tasks they mention, what about the risks of reinforcing biases by using a model's output to train new models? Good grief this paper is lazy!

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#184
post #111

Curious here - OpenAI talks a LOT about how RLHF (Reinforcement Learning Through Human Feedback) is core to how GPT is tuned. Including safety. Are we getting to the point where GPT will be tuned by GPT without the need for HF ?

Maybe, others are using ai for this task, the term to look up is RLAIF.

thank you for this! is the Llama/alpaca tuning the same ?

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#185

Earlier quoted context omitted.

GPT6 or whatever will always require alignment, as the base model just blindly predicts next token, instead of being a helpful, chat style assistant. Right now the best way to align it is with RLHF. The specific technique might change, but in the end there will always be at some level some human input that tells it how it should behave. Newer techniques might further leverage LLMs and require fewer human input. Could…

you don't need humans in the loop for alignment. rlaif is a thing and is used for the anthropic models (claude)

is it really being used for the final model ? i know they have research papers out on it...but wasnt sure if the production models used it.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#186
post #179

Earlier quoted context omitted.

> At 0.06 cents per task Do you work at Verizon?

I got that figure from https://news.ycombinator.com/item?id=35335558 [append] Oh, it's the xkcd Verizon agent reference. Got it. My bad.

I was nitpicking about 0.06 dollars versus 0.06 cents. It's this ancient meme, where a guy records a customer service call with Verizon where they have some back-and-forth about whether these are the same or not.

https://www.youtube.com/watch?v=MShv_74FNWU

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#187
post #167

Earlier quoted context omitted.

Your comment is based on a very strong assumption that all human annotators are alike in their motivations and abilities to do the tasks. As someone who has used mTurk in the past quite a lot, I think this assumption is wrong. That was the reason I stopped using mTurk. On a separate note, how do you type those arrows and subscripts in an HN comment?

This is true. I set up a data validation project with mTurk in the past to validate scraped data by presenting a simply survey of the results. Basically "go to this webpage. The title is X (True or False). The description is Y (True or False)" etc. There were several users who would speedrun the surveys which created a lot of false positives mTurk has some ways to make results more accurate, such as being able to pro…

Yes, that's a lot of work. I remember in one image annotation task we got back what looked like a uniform distribution of responses. This was a pilot to determine the quality of responses so we did not waste a lot of money on it. The more limits we use, the smaller is the pool of mTurkers making it impractical for large tasks.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#188

Earlier quoted context omitted.

The "gold standard" in the article is two graduate students. The Mechanical Turks are probably more- I can't find the number in the paper. Given the "inter-coder agreement" is low for the Mechanical Turks (0.17% Pearson cor. coeff) it is no surprise that the "gold standard" and the Mechanical Turks' decisions diverge. So the comparison is pointless.

It's the gold standard of text labeling to use grad students to label data as a gold standard, and it really depends on the difficulty of the task whether this makes sense. For example, if your labeling task is to, say, transcribe a bunch of recordings of spoken English in American English, maybe a grad student is a nice short-hand for "native American English speaker who is reasonably literate" which is a good basel…

>> Side-notes: - I looked for the 0.17% number in the paper, and what the paper actually says is that the Pearson correlation coefficient between inter-coder agreement and classification accuracy is 0.17 (=positive but weak). This isn't a number comparing the correlation between different coders. (Frankly, I think comparing ChatGPT's self-consistency to consistency between different people is unenlightening) - Section 4 sub-section "Crowd-Workers Annotation" says that each tweet was classified by two different workers, and no single worker classified more than 20% of the dataset.

You're right, I got confused. Quoting from the paper:

>> The relationship between intercoder agreement and accuracy is positive, but weak (Pearson’s correlation coefficient: 0.17).

I thought they used agreement between Mechanical Turks to evaluate their annotations. Thanks for the kind correction.

Regarding grad students, well, most grad students at my (UK) university (Imperial College) are mainly not native English speakers. But that's besides the point I think. My problem is that they took two grad students and compared them to a number of Mechanical Turks. That's just asking for a difference in disagreement between the two groups of annotators, especially if the grad students could communicate with each other and the Mechanical Turks not (which I suspect was the case).

Anyway that's a different criticism than the one in my post above, but I think also useful to keep in mind.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#189
post #145

Suppose you have two classifiers, A and B, and some un-annotated data, D. You want to know how good is classifier B at annotating the data, compared to classifier A. One problem is that you don't have the ground truth for D. So you start by annotating D with the labels assinged by a third classifier, C: C(D) → D₁ Having thus established a modicum of "ground truth", ish, you proceed to annotate D with the two classifi…

Ah yes, ye olde "all humans are equally capable of all tasks" axiom (the paper is correct and the parent comment is wrong, perhaps obviously)

To be clear, that's a different criticism of the article's methodology than mine, yes? If you assume that the two sets of human annotators are different, then what are you comparing, exactly? The ability of one group to second-guess the other?

There are other issues if you choose to assume that the two groups are fundamentally dissimilar: one group was two grad students, the other a number of Mechanical Turks. You can expect there to be more disagreement between (more than two? I'm not sure) Mechanical Turks than between two grad students (both in political science).

Ultimately, my problem with the study is that the labelling they took as ground truth (i.e. that they compared ChatGPT and Mechanical Turks against) is too uncertain, or even subjective, to know for sure what exactly they found out.

Edit: Oh, wait, I didn't put that in the comment above. Damn. I thought I had.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#190
post #167

Suppose you have two classifiers, A and B, and some un-annotated data, D. You want to know how good is classifier B at annotating the data, compared to classifier A. One problem is that you don't have the ground truth for D. So you start by annotating D with the labels assinged by a third classifier, C: C(D) → D₁ Having thus established a modicum of "ground truth", ish, you proceed to annotate D with the two classifi…

Your comment is based on a very strong assumption that all human annotators are alike in their motivations and abilities to do the tasks. As someone who has used mTurk in the past quite a lot, I think this assumption is wrong. That was the reason I stopped using mTurk. On a separate note, how do you type those arrows and subscripts in an HN comment?

>> On a separate note, how do you type those arrows and subscripts in an HN comment?

I type in gvim. It lets you type "digraphs" with Ctrl + K and another couple of keys. For the right-arrow it's (in insert mode and without typing spaces):

  Ctrl K - >
In gvim, you can see all the digraphs with :digraph (in command mode).

The digraph you want to type has to be supported by the font you use of course.

gvim shows you the available digraphs in a scratch buffer that's not terribly easy to search (there's lots of them) but you can yank it to a register and paste it into a new buffer. I don't remember how I do that, I'd have to look up my vim notes :)

>> Your comment is based on a very strong assumption that all human annotators are alike in their motivations and abilities to do the tasks.

Yes and no. There's a few more threads here that discuss this as if ChatGPT beats all humans in annotating data, more or less. I think that's the general assumption anyway.

Also, the way the study was carried out it seems to me that the authors themselves made this assumption, that one group of human annotators is like the other, and you can measure the difference between them as if one was the ground truth and the other just trying to approximate it, where the reality is that they are probably both trying to approximate some completely subjective measure of goodness.

Post reply on HN