Live data from Hacker News

ChatGPT outperforms crowd-workers for text-annotation tasks

arxiv.org

191–200 of 206 posts

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#191

It should comes with no surprise to anyone. Text classification is an easy task. ChatGPT is definitely overqualified to perform this task.

"Classify the following strings according to whether they represent Python programs which terminate..."

Human can't do it either

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#192

Earlier quoted context omitted.

you don't need humans in the loop for alignment. rlaif is a thing and is used for the anthropic models (claude)

is it really being used for the final model ? i know they have research papers out on it...but wasnt sure if the production models used it.

Yeah it is.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#193
post #86

Earlier quoted context omitted.

i think this would better as a browser extension for those who want it, the simplicity of HN is nice

One benefit of the summary bots on Reddit is that people who don't read the article, but read the summary, tend to reply to the summary. I can mostly ignore those comments as they tend to be lower quality. Unfortunately, the people who only read the headline tend to top post (with their comments being lowest quality).

I think ignoring comments and news stories in general on reddit is the best strategy

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#194
post #180
post #113

Earlier quoted context omitted.

Even before LLMs, MTurk has been in a war with the botters, and my understanding is that the MTurkers have to periodically do captchas but many still use bots or various automated tools or utilities. The LLMs and especially the multimodal ones will only make the bots even better and break the captchas even harder. MTurk's business model specifically wants human workers categorized into fine grained categories so they…

Is there such a thing as a captcha in the age of GPT4? Maybe openAI put safety barriers on their model... But surely the botters have a model without that by now?

If you have to select the matching images from 9 images and you are feeding each of those images to GPT4 and asking "is this a ?", I don't think you're going to get a response back to complete the captcha in time (e.g. <30s). This will change as the models advance.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#195
post #12
post #7

Before you ask, because I was curious, from the paper: "For MTurk, we aimed to select the best available crowd-workers, notably by filtering for workers who are classified as “MTurk Masters” by Amazon, who have an approval rate of over 90%, and who are located in the US."

Also "the per-annotation cost of ChatGPT is less than $0.003 -- about twenty times cheaper than MTurk." It's interesting that the best available MTurk Master crowd-workers located in the US are paid about six cents per task.

The per annotation cost of simple image tasks in MTurk can be as low as $0.0003

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#196

I wonder how much we can extend this to image or video labelling tasks? this could really jeopardize the business moats of lot of data labelling startups and services.

You can extend it to data labelling tasks and it will do pretty well.

https://imgur.com/a/3XfOBq1

This is stable diffusion processing a small, blurry image on my PC and correctly identifying it. It did take 7 minutes but my 1 GPU is no match for the ~250M USD worth of GPUs owned by OpenAI.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#197

Earlier quoted context omitted.

File an order and explains how a human being can get paid to do particular tasks? You’d suppose that it would get caught when it fails to pay, but it might be a challenge to arrest an AI in the wild.

Hyperbole. If we wanted to we can always pull the plug, because we exist in neat space and they don’t.

These scenarios get me thinking, how well could an AGI disguise itself as a corporation? Could they start out as a small business working fully remote running some kind of SaaS and scale over a believable timeline? Once they had secured their brand and income would it be possible for them to contract all of the physical work they needed done? One might expect the physical presence of a person in many cases to sign papers or inspect a job site but they could be employees hired by “management” that they never met in person.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#198
post #167

Earlier quoted context omitted.

Your comment is based on a very strong assumption that all human annotators are alike in their motivations and abilities to do the tasks. As someone who has used mTurk in the past quite a lot, I think this assumption is wrong. That was the reason I stopped using mTurk. On a separate note, how do you type those arrows and subscripts in an HN comment?

>> On a separate note, how do you type those arrows and subscripts in an HN comment? I type in gvim. It lets you type "digraphs" with Ctrl + K and another couple of keys. For the right-arrow it's (in insert mode and without typing spaces): Ctrl K - > In gvim, you can see all the digraphs with :digraph (in command mode). The digraph you want to type has to be supported by the font you use of course. gvim shows you the…

Thanks for explaining how you used gvim for writing the comment!

Btw, the authors may make that assumption but it's still a very strong assumption. On mTurk there is no reliable filter for the expertise the authors are looking for. So they used generic filters.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#199
post #193

Earlier quoted context omitted.

One benefit of the summary bots on Reddit is that people who don't read the article, but read the summary, tend to reply to the summary. I can mostly ignore those comments as they tend to be lower quality. Unfortunately, the people who only read the headline tend to top post (with their comments being lowest quality).

I think ignoring comments and news stories in general on reddit is the best strategy

Depends on the subreddit, but generally extremely true.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#200
post #145

Earlier quoted context omitted.

Ah yes, ye olde "all humans are equally capable of all tasks" axiom (the paper is correct and the parent comment is wrong, perhaps obviously)

To be clear, that's a different criticism of the article's methodology than mine, yes? If you assume that the two sets of human annotators are different, then what are you comparing, exactly? The ability of one group to second-guess the other? There are other issues if you choose to assume that the two groups are fundamentally dissimilar: one group was two grad students, the other a number of Mechanical Turks. You ca…

In the framing of your original comment:

A: Mturkers B: ChatGPT C: Experts

=>

ChatGPT better approximates the labelling of D by experts than mTurkers

Which is a coherent and interesting conclusion.

Edit: also, please forgive the snark in my first response

Post reply on HN