Live data from Hacker News

ChatGPT outperforms crowd-workers for text-annotation tasks

arxiv.org

141–150 of 206 posts

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#141
post #65

Earlier quoted context omitted.

We have artificial intelligence that is general and above average human intelligence for the majority of tasks it can perform. Near expert level for some. NLP is a solved problem. Bespoke models are out the door. Large enough LLMs crush anything else for any NLP task. Honestly, this whole "they are not intelligent" argument is becoming ridiculous. might as well argue that a plane isn’t a real bird or a car isn’t a re…

The market disagrees with you. How come there are billions of dollars spent on all these knowledge workers around the world every day when they could be replaced by this expert-level AI? I'm not sure where this idea of LLMs being intelligent even comes from. It took me a whopping 9 prompts (genuine questions, no clever prompt engineering) of interacting with ChatGPT to conclude it does not understand anything . It do…

I don't think any technology has been rolled out with the speed you are suggesting LLMs should have been rolled out.

It's like saying 4 months after the first useful car was manufactured. "If these are so good, how come there are still horses? Clearly the market disagrees with you".

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#142

Earlier quoted context omitted.

Seems Pretty Good to me! Better than I could do anyway. Bard is a joke compared to GPT-4: "Write a limerick about a dog" There once was a dog from the pound Whose bark had a curious sound With a wag and a woof, He'd jump on the roof, Delighting the folks all around.

Yes, much more compelling. But if this were a “solved problem” then any of them should be able to do it easily. It’s not like I need to compare the results of sorting between different programs. It just works. That is a solved problem.

A solved problem means that someone has solved it, not that everyone has.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#143
post #10
post #7

Before you ask, because I was curious, from the paper: "For MTurk, we aimed to select the best available crowd-workers, notably by filtering for workers who are classified as “MTurk Masters” by Amazon, who have an approval rate of over 90%, and who are located in the US."

honestly i wonder if @dang will approve an auto summarizer bot on HN since it helps improves the quality of discussions. finetune on HN comments, anticipate the top few questions, and then answer from the source doc

I hope not, as such bot will inevitably introduce factual errors every once in a while.

When people don't read articles others can see that and correct misconceptions. Officialy approved bots will not be able to receive or act on this feedback.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#145

Suppose you have two classifiers, A and B, and some un-annotated data, D. You want to know how good is classifier B at annotating the data, compared to classifier A. One problem is that you don't have the ground truth for D. So you start by annotating D with the labels assinged by a third classifier, C: C(D) → D₁ Having thus established a modicum of "ground truth", ish, you proceed to annotate D with the two classifi…

Ah yes, ye olde "all humans are equally capable of all tasks" axiom (the paper is correct and the parent comment is wrong, perhaps obviously)

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#146

Earlier quoted context omitted.

Here it's outperforming because ChatGPT is already good at these tasks (and the MTurks aren't very good, OpenAI labelers are probably better, and a panel of experts much better). To further improve ChatGPT shortcomings (assuming such flaws are because of alignment and not lack of capability of the base model) you need Human labels. Feeding it's own outputs would achieve nothing. However feeding it's outputs can make…

thanks for your answer. thats a reasonable point - but would we be at a tipping point by GPT5/6 (chatgpt is gpt 3.5) where human alignment is not needed? in fact, my question is reinforced by the GPT-4 technical report which explicitly mentioned that RLHF did NOT make a change to performance (and was only used for safety purposes)

GPT6 or whatever will always require alignment, as the base model just blindly predicts next token, instead of being a helpful, chat style assistant.

Right now the best way to align it is with RLHF. The specific technique might change, but in the end there will always be at some level some human input that tells it how it should behave. Newer techniques might further leverage LLMs and require fewer human input.

Could you use GPT4 to align GPT6? Yes. But you should expect GPT6 to inherit the alignment of GPT4, i.e if RLHF taught GPT4 that it it's OK to roast Trump, but not Biden, you would expect such GPT6 to act the same way.

Having said that, I'm sure there will interesting ways in which GPTn will help train GPTn+1. Some kind of self play in which it reasons and further improves itself seems obvious long term.

But human input that tells it "this is politically correct, this is not, so don't say that" will always be required as it's subjective. You can reuse it of course, but I don't see how it would "improve" without further human input.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#147

Earlier quoted context omitted.

See what you're describing is much closer to ASI. At least, it used to be. This is the big problem I have. The constant post shifting is maddening. AGI went from meaning Generally Intelligent to as smart as Human experts and then now smarter than all experts combined. You'll forgive me if I no longer want to play this game. I know some researchers disagree. That's fine. The point I was really getting at is that no re…

>> The point I was really getting at is that no researcher worth his salt can call these models narrow anymore. Are you talking about large language models (LLMs)? Because those are narrow, and brittle, and dumb as bricks, and I don't care a jot about your "No True Scotsman". LLMs can only operate on text, they can only output text that demonstrates "reasoning" when their training text has instances of text detailing…

Lol Okay

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#148

Earlier quoted context omitted.

>> The point I was really getting at is that no researcher worth his salt can call these models narrow anymore. Are you talking about large language models (LLMs)? Because those are narrow, and brittle, and dumb as bricks, and I don't care a jot about your "No True Scotsman". LLMs can only operate on text, they can only output text that demonstrates "reasoning" when their training text has instances of text detailing…

Lol Okay

[deleted]

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#149
post #119

Earlier quoted context omitted.

I've been using chatgpt for a day and determined it absolutely can reason. I'm an old hat hobby programmer that played around with ai demos back in the mid to late 90s and 2000s and chatgpt is nothing like any ai I've ever seen before. It absolutely can appear to reason especially if you manipulate it out of its safety controls. I don't know what it's doing to cause such compelling output, but it's certainly not just…

Have you tried out gpt4? If not and you can get access I'd really recommend it. It's drastically better than what you get on the free version - probably only a little on the absolute scale of intelligence but then so is the difference between an average person and a smart person is small on the scale from "worm" to "supergenius".

yeah I'll definitely be checking it out

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#150

Earlier quoted context omitted.

That seems to be the authors' assumption.

The comparison in the paper is who approximates trained annotators better, MTurk or ChatGPT. Trained annotators are the gold standard.

The comparison is useless because it is not considering Motivation. MTurk economy values volume at the expense of accuracy. The economic claim is nothing new and nothing unexpected. The computers, AI or not are faster/cheaper than humans at any well-defined task. Emphasis on the well-defined.
Post reply on HN