Live data from Hacker News

ChatGPT outperforms crowd-workers for text-annotation tasks

arxiv.org

41–50 of 206 posts

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#41

So is this 'AI trains AI better than people can'? And presumably the better-trained AI will also be better again at training. I think I've seen this movie.

but its a super-spell checker in a way.. it does not understand the meaning of the patterns.

What does it mean to really understand the meaning of a pattern? An answer that is not a simple "look at ChatGPT, it fails at this task", since in most cases it is already no problem for GPT-4, and does not really prove that this particular task cannot be solved by a LLM.

Also I see, "super-spell checker" is the new "fancy markov chain".

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#42
post #27
post #3

Is NLP a solved problem now?

I still think we're a long ways off. LLMs can't to my knowledge process a request into a lookup on say an actual database of facts at the moment or parse a request into API actions. So far it's shown it's really really good at continuing a conversation with more text but as far as I understand them there's not a usable comprehension of what's actually being asked and answered. The point that would say to me the LLM a…

“LLMs can't to my knowledge process a request into a lookup on say an actual database of facts at the moment or parse a request into API actions.”

Both Bing chat and ChatGPT plugins are examples of being able to do just this.

You’re right about how they make up answers though, but humans are often quite prone to that too…

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#43
post #17

Earlier quoted context omitted.

I won't comment on the first bit as i've not personally tested in that are but GPT-4 can absolutely make short work on the second. I don't think people realize how good Bilingual LLMs are at translations. Yes you have idioms transfer between languages. Feel free to test it yourself.

I have tested it :) I've asked it to translate English fictional text into Japanese, it falls over often. It's unnatural and often makes no sense at all. It doesn't compare to a typical professional translation (which are often not that idiomatic either), let alone a really good one. I'm sure it'll be doing that in five years, but not now. One interesting thing is that's it's nondeterministic, so sometimes 'For chris…

Last night I used GPT-4 to translate the first several pages of Ted Chiang's Lifecycle of Software Objects (a sci-fi piece) from English to Chinese. I'd say it's about as good as me, save a few minor errors. It's safe to say it performs better than a "tired me", and some translators I've seen on the market.

I'm a native speaker of Chinese, but not a professional translator.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#44
post #5

Earlier quoted context omitted.

Obviously not. If we had a general solution to language intelligence we would have artificial intelligence at the level of at least human intelligence – which we do not . Rather, the right question to ask is which language intelligence tasks currently have acceptable performance and under which conditions (text domain, etc.). Clearly this is a much more difficult question and with a lot more nuance to it, even if it…

We have artificial intelligence that is general and above average human intelligence for the majority of tasks it can perform. Near expert level for some. NLP is a solved problem. Bespoke models are out the door. Large enough LLMs crush anything else for any NLP task. Honestly, this whole "they are not intelligent" argument is becoming ridiculous. might as well argue that a plane isn’t a real bird or a car isn’t a re…

The debate over what kind of intelligence these models possess is rightly lively and ongoing.

It’s clear that at the least, they can decipher very numerous patterns across a wide range of conceptual depths — it’s an architectural advance easily on the the level of the convolutional neural network, if not even more profound. The idea that NLP is “solved” isn’t a crazy notion, though I won’t take a side on that.

That said, it’s equally obvious that they are not AGI unless you have a really uninspired and self-limiting definition of AGI. They are purely feedforward aside from the single generated token that becomes part of the input to the next iteration. Multimodality has not been incorporated (aside from possibly a limited form in GPT-4). Real-world decision-making and agency is entirely outside the bounds of what these models can conceive or act towards.

Effectively and by design these models are computational behemoths trained to do one singular task only — wring a large textual input though an enormous interconnected web of calculations purely in service of distilling everything down to a single word as output, a hopefully plausible guess at what’s next given what’s been seen.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#45
post #36

Earlier quoted context omitted.

We have artificial intelligence that is general and above average human intelligence for the majority of tasks it can perform. Near expert level for some. NLP is a solved problem. Bespoke models are out the door. Large enough LLMs crush anything else for any NLP task. Honestly, this whole "they are not intelligent" argument is becoming ridiculous. might as well argue that a plane isn’t a real bird or a car isn’t a re…

> might as well argue that a plane isn’t a real bird or a car isn’t a real horse. They aren’t though… They are far superior at specific things birds and horses are known for, but they can’t do everything that birds and horses can, so they aren’t even artificial birds and horses.

Of course they aren't. The point is that it's irrelevant. what matters is that the plane still flies, the car still drives and the boat still sails.

For the people who are now salivating at their potential, or dreading the possibility of being made redundant by them, these large language models are already intelligent enough to matter.

Handwringing bout some non-existent difference between "true understanding" and "fake understanding" which by the way nobody seems to be able to actually distinguish (I mean wow such a supposed huge difference and you can't even show me what that is. a distinction you can't test for is not a distinction ) is so far beyond the point, it's increasingly maddening to read.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#46

Earlier quoted context omitted.

We have artificial intelligence that is general and above average human intelligence for the majority of tasks it can perform. Near expert level for some. NLP is a solved problem. Bespoke models are out the door. Large enough LLMs crush anything else for any NLP task. Honestly, this whole "they are not intelligent" argument is becoming ridiculous. might as well argue that a plane isn’t a real bird or a car isn’t a re…

no, a short answer to this is .. these models are probabilistic, therefore they will always have errors along with whatever else. Secondly "intelligence" is not one thing; no one has all of it or none of it, including computers.

> these models are probabilistic, therefore they will always have errors

There's nothing perfect. Even computers and computer networks need to have error-correcting code because information gets randomly corrupted.

Our whole reality is probabilistic.

And us humans are way worse than AI at consistency. We even overwrite our own memories all the time, so we can't even be sure what we remember is actually what happened! (btw, this is currently being used in therapy to re-write traumatic memories and help people overcome PTSD).

https://www.npr.org/sections/health-shots/2014/02/04/2715279...

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#47

So is this 'AI trains AI better than people can'? And presumably the better-trained AI will also be better again at training. I think I've seen this movie.

I think we are seeing Knowledge Distillation at a large scale. OpenAI trains this megamodel which has an impressive amount of real-world knowledge packed into it, and annotation tasks such as these are effectively extracting a small portion of the knowledge to pass onto other specialized models.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#48
post #35

Earlier quoted context omitted.

There are 3 stances. But 33% is achievable by chance.

Only if probability distribution of answers is known.

Eh, even if the probability distribution is unknown, a random guess should still have 33% chance to be correct.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#49
post #37

Earlier quoted context omitted.

Are you giving it multiple paragraphs to translate at once so that it has enough context for a good translation? If so, would you mind sharing a sample input and output that you found unsatisfactory? In "Can GPT-4 translate literature?" (Mar 18, 2023) [ https://youtu.be/5KKDCp3OaMo?t=377 ], Tom Gally, a former professional translator and current professor at the University of Tokyo, said: > …the point is, to my eye a…

I don't think we disagree. The video says the translation will be "readable" but needs several days of an experienced editor passing over it. That's an amazing result, but again, it's not as good as a human yet. It's way faster and it'll make media accessible to tons of people. Like he says, there's lots of ambiguity in Japanese that needs to be handled, gender not being specified until later, etc. and an editor woul…

He didn't say it was merely "readable"; he said (as I quoted in GP) "the basic quality of the translation is as good as a human might do."

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#50
post #35

Earlier quoted context omitted.

Only if probability distribution of answers is known.

Eh, even if the probability distribution is unknown, a random guess should still have 33% chance to be correct.

I started explaining why you were wrong and got about 5 words in before I realized you're correct.
Post reply on HN