Live data from Hacker News

ChatGPT outperforms crowd-workers for text-annotation tasks

arxiv.org

71–80 of 206 posts

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#72

Earlier quoted context omitted.

AGI is Artificial General Intelligence. We have absolutely passed the bar of artificial and generally intelligent. It's not my fault goal post shifting is rampant in this field. And you want to know the crazier thing? Evidently a lot of researchers feel similarly too. General Purpose Technologies ( from the Jobs Paper), General Artificial Intelligence (from the creativity paper). Want to know the original title of th…

> ... these large language models are already intelligent enough to matter. I'm definitely not contesting that. I've always considered the idea of "AGI" to mean something of the holy grail of machine learning -- the point at which there is no real point in pursuing further advances in artificial intelligence because the AI itself will discover and apply such augmentations using its own capabilities. I have seen no ev…

See what you're describing is much closer to ASI. At least, it used to be. This is the big problem I have. The constant post shifting is maddening.

AGI went from meaning Generally Intelligent to as smart as Human experts and then now smarter than all experts combined. You'll forgive me if I no longer want to play this game.

I know some researchers disagree. That's fine. The point I was really getting at is that no researcher worth his salt can call these models narrow anymore. There's absolutely nothing narrow about GPT and the like. So if you think it's not AGI, you've come to accept it no longer means general intelligence.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#73
post #42
post #27

Earlier quoted context omitted.

I still think we're a long ways off. LLMs can't to my knowledge process a request into a lookup on say an actual database of facts at the moment or parse a request into API actions. So far it's shown it's really really good at continuing a conversation with more text but as far as I understand them there's not a usable comprehension of what's actually being asked and answered. The point that would say to me the LLM a…

“LLMs can't to my knowledge process a request into a lookup on say an actual database of facts at the moment or parse a request into API actions.” Both Bing chat and ChatGPT plugins are examples of being able to do just this. You’re right about how they make up answers though, but humans are often quite prone to that too…

A human, if not incentivized to lie or directly incentivized to be truthful, could at least tell you when they're making something up themselves where Bing/Bard seemingly cannot. Once it can do that I think they'll be far more useful, at least then you can have a rough idea of how much you need to check the bots work. If I have to do that for every thing it spits out the best it can do for me is give me new words to use while searching.

Granted getting the name for something to search is often half the battle in tech.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#74
post #20

My main take away here is that Turkers are terrible at some of these tasks. The "stance" task is, "Classify the tweet as having a positive stance towards Section 230, a negative stance, or a neutral stance.", and the Turkers accuracy was like 20%. Even in its best task, ChatGPT only got 75% accuracy.

A cynic might say that someone who could do that task well has better things to do with their time.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#75

Earlier quoted context omitted.

Oh for sure i don't mean to say it's excellent in every language. But i personally think a lot of that is training data representation. Doesn't need to be anywhere equal but for instance after English(93%), the biggest language representation for GPT-3's training corpus is french at...1.8 % Pretty wild. Of Course i don't know the data for GPT-4

I am sure it will improve even further as you pointed out the languages outside of English are fairly low in data represented. However, I guess you said you speak Chinese correct? How well does it do with certain things like older poetic Chinese hanzi? In Japanese if there is a string of kanji it tends to mess up the context. Another area of Japanese it seems poorest at is keigo or polite business Japanese. The way y…

I didn't try with old poetic stuff. Passages sampled from 5 books released in the last 2 decades. You can see what I did thoroughly here. Before GPT-4. Basically a comparison between GLM-130b (English/Chinese model) vs Deepl, Google chatGPT(3.5) etc https://github.com/ogkalu2/Human-parity-on-machine-translati...

Mandarin isn't the second language I speak but I officially compared with it because I wanted to test also with a model that had more equivalent corpus training than the very lopsided gpt models. And Chinese/English is the only combo that has a model of note in that regard.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#76
post #65

Earlier quoted context omitted.

We have artificial intelligence that is general and above average human intelligence for the majority of tasks it can perform. Near expert level for some. NLP is a solved problem. Bespoke models are out the door. Large enough LLMs crush anything else for any NLP task. Honestly, this whole "they are not intelligent" argument is becoming ridiculous. might as well argue that a plane isn’t a real bird or a car isn’t a re…

The market disagrees with you. How come there are billions of dollars spent on all these knowledge workers around the world every day when they could be replaced by this expert-level AI? I'm not sure where this idea of LLMs being intelligent even comes from. It took me a whopping 9 prompts (genuine questions, no clever prompt engineering) of interacting with ChatGPT to conclude it does not understand anything . It do…

To give an example of the limitations of these things that's hopefully easy to understand, I got access to Bard this morning and asked it to write a limerick. It gave me what could charitably be called a free verse poem that happened to begin "there once was a man from Nantucket." I'm sure they can improve on it (ChatGPT was better at this kind of thing when I had access to it) but "solved problem" is clearly a long way off.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#77
post #17

Earlier quoted context omitted.

I won't comment on the first bit as i've not personally tested in that are but GPT-4 can absolutely make short work on the second. I don't think people realize how good Bilingual LLMs are at translations. Yes you have idioms transfer between languages. Feel free to test it yourself.

I have tested it :) I've asked it to translate English fictional text into Japanese, it falls over often. It's unnatural and often makes no sense at all. It doesn't compare to a typical professional translation (which are often not that idiomatic either), let alone a really good one. I'm sure it'll be doing that in five years, but not now. One interesting thing is that's it's nondeterministic, so sometimes 'For chris…

Here is GPT-4's translation, and I find no issues: https://imgur.com/a/oOtf4RD

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#78

Earlier quoted context omitted.

Call me when server farms can reproduce on their own, and network together, and source electricity.

It's not impossible. They could start directing humans to do these thing. I'll keep my phone ready.

> It's not impossible. They could start directing humans to do these thing.

People are already doing this.

Lookup HustleGPT. People are making ChatGPT the CEO/decision maker of their companies/businesses.

There’s an e-commerce that launched 10 days ago, has generated millions of views, over $10k in sales, and the CEO is ChatGPT (https://www.linkedin.com/posts/joao-ferrao-dos-santos_10k-ec...)

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#79
post #38

Earlier quoted context omitted.

That's exactly what's coming. AI will train AI. At some point AI is going to stop needing us to keep evolving. Not sure what that will be like.

Call me when server farms can reproduce on their own, and network together, and source electricity.

Crypto is actually the solution to that. Unlike traditional finance, you don't need a human to sign up under an account. So an AI can just keep it's own wallet and order humans to set up server farms.

Re: ChatGPT outperforms crowd-workers for text-annotation tasks

#80
post #55

In my experience GPT labeling is good, but it makes errors, about 10% in my case. Maybe it's better on average than a human, but not perfect for the task. I was doing open-ended schema matching, a hard task because of the vast number of filed names.

Did you try using multiple AIs for the same task (e.g. GPT & llama), or GPT with two-three prompts?
Post reply on HN