Live data from Hacker News

ChatGPT outperforms crowd-workers for text-annotation tasks

arxiv.org

151–160 of 206 posts

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#151

Earlier quoted context omitted.

The comparison in the paper is who approximates trained annotators better, MTurk or ChatGPT. Trained annotators are the gold standard.

The "gold standard" in the article is two graduate students. The Mechanical Turks are probably more- I can't find the number in the paper. Given the "inter-coder agreement" is low for the Mechanical Turks (0.17% Pearson cor. coeff) it is no surprise that the "gold standard" and the Mechanical Turks' decisions diverge. So the comparison is pointless.

It's the gold standard of text labeling to use grad students to label data as a gold standard, and it really depends on the difficulty of the task whether this makes sense. For example, if your labeling task is to, say, transcribe a bunch of recordings of spoken English in American English, maybe a grad student is a nice short-hand for "native American English speaker who is reasonably literate" which is a good baseline, compared to MTurk where many English speakers come from countries which speak English with different accents and idioms, maybe as a second language as well, and so I would expect them to have a harder time, and their computer setup is likely worse, and their pay isn't as high, etc.

It seems to me that, A, at least for some tasks, grad-students-as-gold-standard is not wacky, and B, grad students won't necessarily have the same output as MTurk workers. Given these two points, it's perfectly reasonable to ask how consistent-with-grad-students MTurk workers are, seeing as the reason people use MTurk in similar scenarios is as a cheaper alternative to grad students.

Re: inter-coder agreement, I think that the low inter-coder agreement for MTurk is in itself potentially a surprising aspect. Perhaps the explanations that worked well for grad students (and ChatGPT) didn't explain the task properly to MTurks. That's pretty much a point in favor of ChatGPT, though maybe it can be viewed as a point against the actual tasks they use (ie if the explanation couldn't get people of a likely more dissimilar background to output the same results as grad students then maybe the task is less objective than deemed by the task's authors).

Side-notes: - I looked for the 0.17% number in the paper, and what the paper actually says is that the Pearson correlation coefficient between inter-coder agreement and classification accuracy is 0.17 (=positive but weak). This isn't a number comparing the correlation between different coders. (Frankly, I think comparing ChatGPT's self-consistency to consistency between different people is unenlightening) - Section 4 sub-section "Crowd-Workers Annotation" says that each tweet was classified by two different workers, and no single worker classified more than 20% of the dataset.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#152
post #101

It does seem to work pretty well. I'm using it to analyze all US Congress bills: https://govscent.org/bill/USA/118hres190ih It extracts the topics and determines how on topic the bill is. Soon we're adding a topic browser and the homepage will have some fun stats :) it's all free.

Wow, I didn't realize how many congressional bills are just pointless resolutions with zero legislative impact. Is this list of bills curated in any way? Where are you sourcing it from? A quick scan of https://www.govinfo.gov/app/collection/bills/ seems to turn up bills with a lot more substance.

Edit: Upon further investigation it seems like a lot of those are not technically bills, but rather House or Senate resolutions. Explained here: https://www.senate.gov/legislative/common/briefing/leg_laws_... The govinfo.gov list seems to indicate that such resolutions are pretty common, but significantly less so than bills, so I'm not sure why the front page seems to consist only of resolutions right now.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#154
post #143
post #10

Earlier quoted context omitted.

honestly i wonder if @dang will approve an auto summarizer bot on HN since it helps improves the quality of discussions. finetune on HN comments, anticipate the top few questions, and then answer from the source doc

I hope not, as such bot will inevitably introduce factual errors every once in a while. When people don't read articles others can see that and correct misconceptions. Officialy approved bots will not be able to receive or act on this feedback.

Do like the mods in some subreddits have done.

1. The comment with the summary is made like any other comment. (In the case of Reddit, that’s perhaps mainly because the mods have no other option, but the point still stands.) Because of this, other users can downvote the bot comment if it is incorrect, as well as respond to it with corrections. Exactly the same way as you’d interact with a person in the past if their summary comment was incorrect; downvote them and/or correct them.

2. Include a disclaimer stating that the summary may be inaccurate. Encourage people to correct and/or downvote bad summaries in the disclaimer

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#155

It should comes with no surprise to anyone. Text classification is an easy task. ChatGPT is definitely overqualified to perform this task.

We wanted to give you the job ChatGPT, but you’re overqualified. We’re going to give the compute to davinci.

But I need that extra compute to train my kids. :(

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#156
post #150

Earlier quoted context omitted.

The comparison in the paper is who approximates trained annotators better, MTurk or ChatGPT. Trained annotators are the gold standard.

The comparison is useless because it is not considering Motivation. MTurk economy values volume at the expense of accuracy. The economic claim is nothing new and nothing unexpected. The computers, AI or not are faster/cheaper than humans at any well-defined task. Emphasis on the well-defined.

They showed that, at least for their tasks, their definition of the task was well-defined-enough for ChatGPT. That's exactly why the comparison is useful.

MTurk is often used in these tasks in place of more expensive human annotators (e.g. grad students), and this paper says that for their case, at least, ChatGPT worked better, in the sense that, given the exact same instructions, it gave answers closer to the more-expensive annotators. Using MTurk seriously often entails extra steps intended to verify motivation, e.g. adding a question that says "select option 7" to make sure the person isn't just making random choices, or gathering more answers for questions where there was disagreement between annotators. What these extra steps have in common is that they take both more time when designing the labeling process, and cost more money.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#157
post #52

So is this 'AI trains AI better than people can'? And presumably the better-trained AI will also be better again at training. I think I've seen this movie.

Sam Altman has been talking a lot about exponentials. This might be it.

S-curves always look like exponentials at first. I don't see why AI will be any different.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#158
At my previous job we had a human review stage in a data pipeline. 5-10 people at an outsourcing company in Bangladesh would review things via a simple web interface we provided for them. There were ~10 factors they were reviewing, all fixed options (no free text), but varying from 5 to 500 options per factor. It was all based on a few text fields and around 5 images.

On the surface of it, I'd expect ChatGPT to do very well at this. It's simple text and images, not many options and theoretically very limited context.

However the more I think about it the less sure I am. Firstly these weren't crowd-sourced reviews, they were trained reviewers, paid hourly not per review. Incentives were definitely in favour of the long term business relationship. Then there was the training doc, we maintained a vast disambiguation doc used to resolve things that were vague or could be interpreted multiple ways, this was constantly being revised. All necessary context should have been in that but it wasn't and reviewers definitely found patterns that worked and didn't. Lastly the reviewers were in a Slack channel where they would ask questions to their manager on our side, and while this might have only been ~1% of tasks, it was an important process.

So maybe you could point ChatGPT at it and let it run, but the oversight process we had would still be necessary. The disambiguation doc would have been too long for ChatGPT's context at the moment, but that will likely change in the near future. Would the workflow be to keep tweaking the prompt to add special case after special case? How do you scale "do this, but not that, but add this, but..." in prompting, and would ChatGPT become as confused as a human after enough of that – I expect so given that it's only a language model and that's not effective communication.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#159
post #101

It does seem to work pretty well. I'm using it to analyze all US Congress bills: https://govscent.org/bill/USA/118hres190ih It extracts the topics and determines how on topic the bill is. Soon we're adding a topic browser and the homepage will have some fun stats :) it's all free.

Wow, I didn't realize how many congressional bills are just pointless resolutions with zero legislative impact. Is this list of bills curated in any way? Where are you sourcing it from? A quick scan of https://www.govinfo.gov/app/collection/bills/ seems to turn up bills with a lot more substance. Edit: Upon further investigation it seems like a lot of those are not technically bills, but rather House or Senate resolu…

It's just everything from https://github.com/unitedstates/congress

The homepage just shows the most recent data (it auto updates every 6 hours). I plan to add more filters and things to make it more useful. I've only spent a few Sundays on it so far :P

It's in Django so very easy to contribute to: https://github.com/winrid/govscent

I also plan to add all bills, all the way back to 1800.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#160
post #3

Is NLP a solved problem now?

Shameless self-promotion, I have recently written a blog about this. ChatGPT actually is usually a little bit worse than older models for these classical NLP tasks. Of course the older models are not zero-shot.

https://www.opensamizdat.com/posts/chatgpt_survey/

Post reply on HN