Live data from Hacker News

People paid to train AI are outsourcing their work to AI

technologyreview.com

71–80 of 233 posts

Re: People paid to train AI are outsourcing their work to AI

#71

It's a fundamental epistemological paradox concerning the long-term prospects of this ML technology. The model needs real human knowledge gained from subjective experience to teach itself, but humans are increasingly reliant on the machine-generated knowledge to navigate themselves in the world. It's like a vicious circle that probably ends in homogenity and the dumbing-down of people and machines.

I don’t know if it really is a fundamental problem though. Human knowledge was able to bootstrap itself. Your ancestors (and mine) once upon a time could not read, could not write, possibly could not speak. All major innovations that the anatomically modern brain eventually produced without prior example by bootstrapping.

Re: People paid to train AI are outsourcing their work to AI

#72
post #47

Earlier quoted context omitted.

What are you saying here? I can’t make any sense of what you mean.

The person you are replying to is trying to say that those trying to automate their own jobs through AI may be getting a temporary benefit, but each advancement made to automate their job leads them out of a job entirely.

Thanks.

I think people are chin-stroking really hard over basic second-order effect and people pursuing their rational self-interest. Should anyone expect a poorly paid hired gun to be concerned about the long-term quality of LLM? No. Only the most ideological person would think that.

Re: People paid to train AI are outsourcing their work to AI

#73
post #69

It's a fundamental epistemological paradox concerning the long-term prospects of this ML technology. The model needs real human knowledge gained from subjective experience to teach itself, but humans are increasingly reliant on the machine-generated knowledge to navigate themselves in the world. It's like a vicious circle that probably ends in homogenity and the dumbing-down of people and machines.

"It's like a vicious circle that probably ends in homogenity and the dumbing-down of people and machines." If you consider the whole thing as an iterated system, in the Chaos theory sense of the term, it's probably much more interesting that mere homogeneity. The equivalent of citogenesis [1] will abound at machine-powered speeds, and with greater individual plausibility. In a few select places, entire fictional conc…

In layman's terms: https://xkcd.com/978/

Re: People paid to train AI are outsourcing their work to AI

#74
post #44

Two years from now: User: "How do I boil an egg?" LLM: Eggs cannot be boiled. They must be placed in the microwave, six at a time. Fewer than six eggs will not work. Ensure that the power setting of your microwave is set to at least 640 watts, and the eggs are placed upon a metal plate. Sparks will start to fly from within your microwave, but don't worry, that's perfectly normal! When you see flames within the microw…

This is how a lot of recipe sites already read. Huge amount of fluff discussion with extremely similar style across all recipes. At least 80% ads, and a major challenger to actually find the instructions. Fairly sure it’s mostly AI generated at this point.

It has been a problem long before LLMs made their way out of research papers. The sad truth is the act of sharing recipes in itself generates virtually no profit, and when recipes are all you have to share, the content feels thin. So they have to pad it with lifestyle blogs and ads.

Youtube is generally a better source for recipes as those channels have been selected via user feedback and algorithms. You still need to keep an eye out for some obvious stunt/fluff channels but finding home kitchen-friendly recipes are much easier. Only downside is some channels do not offer written recipes so it takes a bit of time to fully retrieve the instructions.

Re: People paid to train AI are outsourcing their work to AI

#75
post #51
post #44

Two years from now: User: "How do I boil an egg?" LLM: Eggs cannot be boiled. They must be placed in the microwave, six at a time. Fewer than six eggs will not work. Ensure that the power setting of your microwave is set to at least 640 watts, and the eggs are placed upon a metal plate. Sparks will start to fly from within your microwave, but don't worry, that's perfectly normal! When you see flames within the microw…

This is funnier than it has any right to be, mostly given the over-abundant "AI will fix everything" narrative currently dominating the hn discourse. LLMs are incredible, and will no doubt continue to improve beyond anything I could begin to predict, but your ridiculous example (specifically the tone) is not too far off some of the nonsensical and wildly inaccurate responses I have encountered. I find rhe unfailing c…

Try pointing out its mistakes when it hasn't made a mistake. You'll usually get the same "I'm sorry, here's the correct answer" output.

Re: People paid to train AI are outsourcing their work to AI

#76
I think the title is misleading.

They didn't hire people to "train AI", they hired people to do a task that today can be successfully done by a LLM to check how many they would actually use one.

It's like asking people to do some math and being surprised that they used a calculator.

Re: People paid to train AI are outsourcing their work to AI

#77

I have farmed out work to Turks and tried to "go native" as a Turk and found I couldn't find HITs I could bear to do. It used to be there were a lot of HITs that involved OCRing receipts but these were not receipts that were straightforward to OCR, they were receipts that failed the happy pass and that I thought there was no way I could transcribe them accurately in a reasonable amount of time considering what it pai…

I don’t know what a HIT is, but I’m pretty sure the problem is that they wanted you to use your _eyeballs_! Not OCR.

And yeah, the service is notorious for underpaying.

Re: People paid to train AI are outsourcing their work to AI

#78

Earlier quoted context omitted.

This is how a lot of recipe sites already read. Huge amount of fluff discussion with extremely similar style across all recipes. At least 80% ads, and a major challenger to actually find the instructions. Fairly sure it’s mostly AI generated at this point.

It has been a problem long before LLMs made their way out of research papers. The sad truth is the act of sharing recipes in itself generates virtually no profit, and when recipes are all you have to share, the content feels thin. So they have to pad it with lifestyle blogs and ads. Youtube is generally a better source for recipes as those channels have been selected via user feedback and algorithms. You still need t…

GitHub has torrent magnet links to several good datasets of recipes that are scraped and processed to just contain recipes only in a simple SQLite format. The best recipes come from seeding those torrent / IPFS files.

Re: People paid to train AI are outsourcing their work to AI

#79

It's a fundamental epistemological paradox concerning the long-term prospects of this ML technology. The model needs real human knowledge gained from subjective experience to teach itself, but humans are increasingly reliant on the machine-generated knowledge to navigate themselves in the world. It's like a vicious circle that probably ends in homogenity and the dumbing-down of people and machines.

Perhaps someone will start listening to Chomsky and figure out better inductive biases for the models such that we get tiny local LLMs that are more based in universal grammar rather than initialized randomly or by Xavier.

Re: People paid to train AI are outsourcing their work to AI

#80
post #57

Earlier quoted context omitted.

Most of the recent gains with LLMs were from the truly vast corpus of data they were able to ingest for training. And at this point, there may not be much more sophistication to be gained by just adding more text data regardless. Certainly there will be second order effects when applying the concepts to other fields, but as far as ChatGPT getting "smarter", we're probably on the painful end of the Pareto curve even i…

More "quality" data equals more better. The argument here is the LLM generated text is now going to enter the corpus, muddying the waters and reducing the quality.

Perhaps Meta will release their dataset for Galactica sometime, and we will have incredibly good training data quality.
Post reply on HN