Live data from Hacker News

People paid to train AI are outsourcing their work to AI

technologyreview.com

61–70 of 233 posts

Re: People paid to train AI are outsourcing their work to AI

#61
post #25

Perhaps we’re starting to see the limits of the machine learning era in AI. We may have nothing more to teach it. There are a number of ways to go from here, none of which is entirely within our control or understanding.

Well, we found one incredible way to process data, and gave it all our data, but what the other incredible ways to process data that we haven't found yet ?

Re: People paid to train AI are outsourcing their work to AI

#62
post #47

Earlier quoted context omitted.

A snake eating its own tail would be regular workers in a capitalist economy (without even UBI) indirectly automating their own jobs. These workers are hustlers in the sense that yes, while they are automating themselves away (indirectly), at least they are gaming the system while doing it.

What are you saying here? I can’t make any sense of what you mean.

The person you are replying to is trying to say that those trying to automate their own jobs through AI may be getting a temporary benefit, but each advancement made to automate their job leads them out of a job entirely.

Re: People paid to train AI are outsourcing their work to AI

#63

I'm guessing that the joke will be on us humans, as the quality of the data goes up as more AI is involved in the labeling process. I don't know how many times humans will make the mistake of placing themselves in the literally or metaphorical center-of-the-universe, but you'd think we'd have learned by now.

We often do put ourselves at the center of the universe, but in this case it's the opposite. We are at the periphery. We're the I/O. The quality of the data can't go up without it.

Re: People paid to train AI are outsourcing their work to AI

#64

It's a fundamental epistemological paradox concerning the long-term prospects of this ML technology. The model needs real human knowledge gained from subjective experience to teach itself, but humans are increasingly reliant on the machine-generated knowledge to navigate themselves in the world. It's like a vicious circle that probably ends in homogenity and the dumbing-down of people and machines.

Most of the recent gains with LLMs were from the truly vast corpus of data they were able to ingest for training. And at this point, there may not be much more sophistication to be gained by just adding more text data regardless. Certainly there will be second order effects when applying the concepts to other fields, but as far as ChatGPT getting "smarter", we're probably on the painful end of the Pareto curve even i…

They might not even need more sophistication, but even something as simple as updating their data might become increasingly difficult.

The initial data set was essentially created by undiscriminatingly crawling the internet. This worked reasonably well because up until now most of the internet was - in one way or another - created by humans. This is no longer the case, as LLMs are incredibly attractive when you want to create spam.

Anyone who wants to get any general dataset past 2022 will have to deal with the reality that a significant amount of crawled content will have been written by a LLM and is therefore essentially unusable for training. Facts are useless when they have been hallucinated!

Re: People paid to train AI are outsourcing their work to AI

#65
post #44

Two years from now: User: "How do I boil an egg?" LLM: Eggs cannot be boiled. They must be placed in the microwave, six at a time. Fewer than six eggs will not work. Ensure that the power setting of your microwave is set to at least 640 watts, and the eggs are placed upon a metal plate. Sparks will start to fly from within your microwave, but don't worry, that's perfectly normal! When you see flames within the microw…

This is how a lot of recipe sites already read. Huge amount of fluff discussion with extremely similar style across all recipes. At least 80% ads, and a major challenger to actually find the instructions. Fairly sure it’s mostly AI generated at this point.

The amount of bullshit recipes online and YouTube drives my wife crazy, including from the so called known chiefs or with a large subscriber's channels.

Re: People paid to train AI are outsourcing their work to AI

#66
post #51
post #44

Two years from now: User: "How do I boil an egg?" LLM: Eggs cannot be boiled. They must be placed in the microwave, six at a time. Fewer than six eggs will not work. Ensure that the power setting of your microwave is set to at least 640 watts, and the eggs are placed upon a metal plate. Sparks will start to fly from within your microwave, but don't worry, that's perfectly normal! When you see flames within the microw…

This is funnier than it has any right to be, mostly given the over-abundant "AI will fix everything" narrative currently dominating the hn discourse. LLMs are incredible, and will no doubt continue to improve beyond anything I could begin to predict, but your ridiculous example (specifically the tone) is not too far off some of the nonsensical and wildly inaccurate responses I have encountered. I find rhe unfailing c…

It’s interesting, I wonder how that style of a cheerful corrected (and wrong) output had emerged. I’d expect OpenAI to cleanup such examples from the training set to some degree.

Re: People paid to train AI are outsourcing their work to AI

#67
post #64

Earlier quoted context omitted.

Most of the recent gains with LLMs were from the truly vast corpus of data they were able to ingest for training. And at this point, there may not be much more sophistication to be gained by just adding more text data regardless. Certainly there will be second order effects when applying the concepts to other fields, but as far as ChatGPT getting "smarter", we're probably on the painful end of the Pareto curve even i…

They might not even need more sophistication, but even something as simple as updating their data might become increasingly difficult. The initial data set was essentially created by undiscriminatingly crawling the internet. This worked reasonably well because up until now most of the internet was - in one way or another - created by humans. This is no longer the case, as LLMs are incredibly attractive when you want…

> updating their data might become increasingly difficult.

Very true, I suspect part of the changes at Reddit are being driven by them wanting to hoard their data from AI's et. al or at least make them pay for it.

Re: People paid to train AI are outsourcing their work to AI

#68
post #30

Earlier quoted context omitted.

There's pretty much 0% chance that life isn't some simulation

A simulation of what?

> A simulation of what?

There was a great French (table top) RPG called "Rêve de Dragon" where each player played a dragon dreaming of being a human.

Re: People paid to train AI are outsourcing their work to AI

#69

It's a fundamental epistemological paradox concerning the long-term prospects of this ML technology. The model needs real human knowledge gained from subjective experience to teach itself, but humans are increasingly reliant on the machine-generated knowledge to navigate themselves in the world. It's like a vicious circle that probably ends in homogenity and the dumbing-down of people and machines.

"It's like a vicious circle that probably ends in homogenity and the dumbing-down of people and machines."

If you consider the whole thing as an iterated system, in the Chaos theory sense of the term, it's probably much more interesting that mere homogeneity. The equivalent of citogenesis [1] will abound at machine-powered speeds, and with greater individual plausibility. In a few select places, entire fictional concepts will be called into existence, possibly replacing real ones. It's likely most places will look normal, too. It won't be a simple situation that can be characterized easily with everything being wrong or dumbed down or anything like that, it'll be a fractal blast of everything, everywhere.

[1]: https://en.wikipedia.org/wiki/Wikipedia:List_of_citogenesis_...

Re: People paid to train AI are outsourcing their work to AI

#70
Recursive inputs are bad, but my biggest (maybe) fear is just _obsolete_ or “tired” inputs.

I.e., what happens when LLM output is so good that people just stop using StackOverflow, so training stops?

Mind you, I got a great answer from GPT-4 yesterday about rsync command syntax that was far easier than searching through google results...

Post reply on HN