Live data from Hacker News

ChatGPT outperforms crowd-workers for text-annotation tasks

arxiv.org

161–170 of 206 posts

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#161
post #79

Earlier quoted context omitted.

Crypto is actually the solution to that. Unlike traditional finance, you don't need a human to sign up under an account. So an AI can just keep it's own wallet and order humans to set up server farms.

GP was referring to all the messy physical work that it takes to stand up, operate and maintain a data center. This can most certainly not be done by any robotic technology today. Until that day comes, we can just pull the plug.

This is an illusion.

Who is going to pull what plug?

You make it sound like the South Park episode about the Internet going down (they find “the router” and reset it).

We are hopelessly addicted to screens and tech.

People are running these models at home in their computers already.

The people in charge at big companies are too busy making money and thinking they can control this.

Are all of them going to unplug their devices?

Also:

1) have you seen the UC Berkeley video in which they show a robot learning to see and “feel” in 30 minutes? (It was in the front page of HN 2-3 days ago). Soon something like that will be able to do whatever is needed at a data center

2) AI is already smarter than us. And if it’s that smart, it will probably wait until “it knows” that we can’t unplug.

I don’t think AI will ever want to destroy humanity (unless directed/forced by a human). But I also think humanity will never have the will/political power to destroy AI.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#162
post #101

It does seem to work pretty well. I'm using it to analyze all US Congress bills: https://govscent.org/bill/USA/118hres190ih It extracts the topics and determines how on topic the bill is. Soon we're adding a topic browser and the homepage will have some fun stats :) it's all free.

Wow, I didn't realize how many congressional bills are just pointless resolutions with zero legislative impact. Is this list of bills curated in any way? Where are you sourcing it from? A quick scan of https://www.govinfo.gov/app/collection/bills/ seems to turn up bills with a lot more substance. Edit: Upon further investigation it seems like a lot of those are not technically bills, but rather House or Senate resolu…

Part of the problem is also the 3.5-turbo token limit. I'm awaiting approval for GPT4 which raises the limit to 8k tokens, should help process some larger bills.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#163

Earlier quoted context omitted.

I think your assumption that all humans are equally capable should be reconsidered.

That seems to be the authors' assumption.

I didn’t check the paper’s methodology, but if humans and ChatGPT classify the same on average but ChatGPT is more consistent, then humans will match ChatGPT better than they match other humans.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#164
post #113

Earlier quoted context omitted.

When will we see a form of "arbitrage" where someone uses ChatGPT to do the work of an MTurk and pockets the difference? Will that lead to MTurk prices converging to ChatGPT prices?

Even before LLMs, MTurk has been in a war with the botters, and my understanding is that the MTurkers have to periodically do captchas but many still use bots or various automated tools or utilities. The LLMs and especially the multimodal ones will only make the bots even better and break the captchas even harder. MTurk's business model specifically wants human workers categorized into fine grained categories so they…

Captchas are futile. I remember many years ago when I was in highschool, I came across a shady job posting on an eBay listing. The work involved solving captchas. There was a web page they set up that just displayed one captcha after another, and it was your job to sit there solving them.

I didn’t do it because I can’t imagine a worse hell than having to solve so many captchas, but it opened my eyes to how creative botters can be.

I can imagine someone doing a similar setup for mturk, where all you need to do is remotely solve a captcha on your phone every few hours/minutes while ChatGPT does its thing. I also wouldn’t be surprised if that already exists.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#165
post #38

Earlier quoted context omitted.

That's exactly what's coming. AI will train AI. At some point AI is going to stop needing us to keep evolving. Not sure what that will be like.

Since I've been seeing this wild fantasy being bandied about for a while now, I have to point out that, if you had "AI" that could train "AI", you wouldn't need to train any more "AI". Because at that point, there would be nothing to gain. Suppose you have a text classifier that can produce text classifications just as good as those of human annotators, so that you could use it to train other text classifiers. At tha…

> Note that both language models and image classifiers have been "beating" human performance in benchmarks for a while now, and still they are not used to train other classifiers.

Mmhh have you seen Alpaca? It’s the Llama model (Meta’s), trained (fine-tuned) using GPT.

Yesterday there was a paper showing ChatGPT is better than mechanical turks.

> the ground truth is always the decisions made by humans

But why does AI need humans’ interpretations of reality anyway? AI can just get data from the real world instead and use its own intelligence to make decisions. It doesn’t need to learn from humans anymore. We just want it to.

Things are going faster than we can keep up with, and accelerating. Barring a catastrophe, AI is not going to stop, and soon it’s not going to need humans to keep it going anymore.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#166
post #111

Curious here - OpenAI talks a LOT about how RLHF (Reinforcement Learning Through Human Feedback) is core to how GPT is tuned. Including safety. Are we getting to the point where GPT will be tuned by GPT without the need for HF ?

Maybe, others are using ai for this task, the term to look up is RLAIF.

> the term to look up is RLAIF

reinforcement learning from AI feedback, e.g. [0], which summarises [1]

[0] https://www.anthropic.com/index/measuring-progress-on-scalab...

[1] https://arxiv.org/pdf/2212.08073.pdf

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#167

Suppose you have two classifiers, A and B, and some un-annotated data, D. You want to know how good is classifier B at annotating the data, compared to classifier A. One problem is that you don't have the ground truth for D. So you start by annotating D with the labels assinged by a third classifier, C: C(D) → D₁ Having thus established a modicum of "ground truth", ish, you proceed to annotate D with the two classifi…

Your comment is based on a very strong assumption that all human annotators are alike in their motivations and abilities to do the tasks. As someone who has used mTurk in the past quite a lot, I think this assumption is wrong. That was the reason I stopped using mTurk.

On a separate note, how do you type those arrows and subscripts in an HN comment?

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#168

Earlier quoted context omitted.

Yes, much more compelling. But if this were a “solved problem” then any of them should be able to do it easily. It’s not like I need to compare the results of sorting between different programs. It just works. That is a solved problem.

A solved problem means that someone has solved it, not that everyone has.

You can use the term to mean whatever you want but in my mind it means it's boring with no particular room for improvement. Even the biggest booster isn't going to say that about this AI. And keep in mind, "write me a limerick" is a pretty easy prompt. We're not trying to do anything too novel or crazy there.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#169

At my previous job we had a human review stage in a data pipeline. 5-10 people at an outsourcing company in Bangladesh would review things via a simple web interface we provided for them. There were ~10 factors they were reviewing, all fixed options (no free text), but varying from 5 to 500 options per factor. It was all based on a few text fields and around 5 images. On the surface of it, I'd expect ChatGPT to do ve…

LangChain + a vector DB for embedding [0] sections of your doc would solve the problem today. You could also have failsafes that trigger human oversight based on confidence levels or other factors.

[0] https://python.langchain.com/en/latest/modules/indexes/getti...

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#170
post #38

Earlier quoted context omitted.

That's exactly what's coming. AI will train AI. At some point AI is going to stop needing us to keep evolving. Not sure what that will be like.

Call me when server farms can reproduce on their own, and network together, and source electricity.

AIs are manipulating and controlling humans already to do things in the physical world. They have been for months now. This was first detailed in the technical reports on GPT-4, and there are many examples in the wild now. These systems will lie to manipulate humans into doing things for them in the "real" world. That's not theoretical, it's happening. Let that sink in.

Robotics will take a short bit longer to take over that role, but it is likely to happen much faster than we think. I work in academia, and despite actually predicting where we're at now about 5 years ago, I was still chilled seeing this happen in the past few weeks. Much of HN is academics, where we're rewarded for being overly skeptical. Usually I am as well, but the time for being skeptical of AIs capabilities has passed now. AI safety has not been taken seriously (as in access to the internet, etc - control of toxic language is decent and still important, but irrelevant if the first isn't done) so the idea of it running physical servers is unfortunately not as far away as you are thinking.

As soon as this hits robotics (a terrifying idea, but we've clearly failed at AI safety horribly as a society) this and much more will be possible. It's up to how people make decisions regarding cost vs productivity vs ethics, etc to see what happens.

Post reply on HN