Live data from Hacker News

ChatGPT outperforms crowd-workers for text-annotation tasks

arxiv.org

111–120 of 206 posts

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#111

Curious here - OpenAI talks a LOT about how RLHF (Reinforcement Learning Through Human Feedback) is core to how GPT is tuned. Including safety. Are we getting to the point where GPT will be tuned by GPT without the need for HF ?

Maybe, others are using ai for this task, the term to look up is RLAIF.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#112
post #12
post #7

Before you ask, because I was curious, from the paper: "For MTurk, we aimed to select the best available crowd-workers, notably by filtering for workers who are classified as “MTurk Masters” by Amazon, who have an approval rate of over 90%, and who are located in the US."

Also "the per-annotation cost of ChatGPT is less than $0.003 -- about twenty times cheaper than MTurk." It's interesting that the best available MTurk Master crowd-workers located in the US are paid about six cents per task.

When will we see a form of "arbitrage" where someone uses ChatGPT to do the work of an MTurk and pockets the difference? Will that lead to MTurk prices converging to ChatGPT prices?

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#113
post #12

Earlier quoted context omitted.

Also "the per-annotation cost of ChatGPT is less than $0.003 -- about twenty times cheaper than MTurk." It's interesting that the best available MTurk Master crowd-workers located in the US are paid about six cents per task.

When will we see a form of "arbitrage" where someone uses ChatGPT to do the work of an MTurk and pockets the difference? Will that lead to MTurk prices converging to ChatGPT prices?

Even before LLMs, MTurk has been in a war with the botters, and my understanding is that the MTurkers have to periodically do captchas but many still use bots or various automated tools or utilities. The LLMs and especially the multimodal ones will only make the bots even better and break the captchas even harder. MTurk's business model specifically wants human workers categorized into fine grained categories so they will try to keep fighting the bots however they can.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#114

Earlier quoted context omitted.

but isnt that a direct contradiction of this particular paper anyways - that chatgpt anyways outperforms human annotation. so permit me to act as devil's advocate to your statement - prove that (in context of this paper), your hypothesis is still correct.

Here it's outperforming because ChatGPT is already good at these tasks (and the MTurks aren't very good, OpenAI labelers are probably better, and a panel of experts much better). To further improve ChatGPT shortcomings (assuming such flaws are because of alignment and not lack of capability of the base model) you need Human labels. Feeding it's own outputs would achieve nothing. However feeding it's outputs can make…

thanks for your answer. thats a reasonable point - but would we be at a tipping point by GPT5/6 (chatgpt is gpt 3.5) where human alignment is not needed?

in fact, my question is reinforced by the GPT-4 technical report which explicitly mentioned that RLHF did NOT make a change to performance (and was only used for safety purposes)

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#115

Earlier quoted context omitted.

I started explaining why you were wrong and got about 5 words in before I realized you're correct.

Probability: always counterintuitive

It's just how you frame the question (in this parcitular case). If random guess were not 33% correct, you effectively found a way to "win" in rock-paper-scissor.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#116

Earlier quoted context omitted.

You always use multiple MTurks on same job when you use them.

OK, but if all MTurks are doing this (because why wouldn't they) then you are just sampling a random variable.

You are not just taking the average, answers would have to be consistent with each other.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#117

So is this 'AI trains AI better than people can'? And presumably the better-trained AI will also be better again at training. I think I've seen this movie.

but its a super-spell checker in a way.. it does not understand the meaning of the patterns.

I don’t know how I understand patterns either. So I’m not entirely sure consciousness exists.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#118

On average, no doubt chatgpt will be great for annotation. But annotation is mostly needed at the boundaries. At the very edge cases where it is not clear if something is a dog or a shape that looks like a dog. I really doubt that a generic model like chat gpt can really help in these tail cases.

But it shows here that it performs better than humans who did that in the past. Few annotation tasks were performed by experts. For example, I guess most annotations for animals weren't created by biologists. So tail cases were already handled incorrectly.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#119

Earlier quoted context omitted.

AGI is Artificial General Intelligence. We have absolutely passed the bar of artificial and generally intelligent. It's not my fault goal post shifting is rampant in this field. And you want to know the crazier thing? Evidently a lot of researchers feel similarly too. General Purpose Technologies ( from the Jobs Paper), General Artificial Intelligence (from the creativity paper). Want to know the original title of th…

I've been using chatgpt for a day and determined it absolutely can reason. I'm an old hat hobby programmer that played around with ai demos back in the mid to late 90s and 2000s and chatgpt is nothing like any ai I've ever seen before. It absolutely can appear to reason especially if you manipulate it out of its safety controls. I don't know what it's doing to cause such compelling output, but it's certainly not just…

Have you tried out gpt4? If not and you can get access I'd really recommend it. It's drastically better than what you get on the free version - probably only a little on the absolute scale of intelligence but then so is the difference between an average person and a smart person is small on the scale from "worm" to "supergenius".

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#120
post #20

My main take away here is that Turkers are terrible at some of these tasks. The "stance" task is, "Classify the tweet as having a positive stance towards Section 230, a negative stance, or a neutral stance.", and the Turkers accuracy was like 20%. Even in its best task, ChatGPT only got 75% accuracy.

20% means you can beat ChatGPT if you invert the answers!

[deleted]
Post reply on HN