Well, there is a very simple solution to that problem: Pay people extra money for the privilege of having them fulfill the tasks the way you want them to.
Or just use the tools directly since they no longer underperform human actors.
This is going to cause problems as many Mechanical Turk tasks are being used to train AI models...
And even fine-tuning using AI causes the models to quickly degrade I've seen here previously. Perhaps we are already at peak of these LLMs now and the next gen won't come because all training data will be poisened with AI generate crap. That would be kinda funny to me tbh
The invention of nukes caused a similar problem. Scientists had to harvest metal from pre-ww2 sunken battleships, since that’s the only metal not contaminated with some measurable amount of radiation. I forget the details, but it’s very likely pre-chatgpt training data will be prioritized, with exceptions for news and current events (which largely doesn’t matter whether it’s AI generated anyway).
This is not novel. Mturk has always been automated since its very first release.
I remember worker forums sharing scripts to query amazon for the expected results that would get you paid. Amazon put up a bunch of very lucrative test jobs, but you only got paid if your answer matched the majority. The majority were automating, so unless you used the sometimes incorrect answers, you got nothing.
Well, there is a very simple solution to that problem: Pay people extra money for the privilege of having them fulfill the tasks the way you want them to.
Which would mean working in a controlled environment or installing invasive software in their computers (something that would block chatgpt for example)
This is going to cause problems as many Mechanical Turk tasks are being used to train AI models...
And even fine-tuning using AI causes the models to quickly degrade I've seen here previously. Perhaps we are already at peak of these LLMs now and the next gen won't come because all training data will be poisened with AI generate crap. That would be kinda funny to me tbh
I have a feeling that some types of data will become a bit like "the DNA wealth of the rainforest", valuable because it's so rich, different and rare.
In particular small and/or dying human languages. Languages carry with them an immense amount of information, from all the precious human experiences that have shaped them.
Well, there is a very simple solution to that problem: Pay people extra money for the privilege of having them fulfill the tasks the way you want them to.
Or just use the tools directly since they no longer underperform human actors. GPT-4 going toe to toe with the experts who set the benchmarks, way overperforming the crowdworkers https://www.artisana.ai/articles/gpt-4-outperforms-elite-cro... GPT-3.5 outperforming crowdworkers https://arxiv.org/abs/2303.15056 More money won't stop people from automating whatever they can. Especially when the automation provides such…
How did the researchers work out which of the human or GPT4 response was the correct one?
Well, there is a very simple solution to that problem: Pay people extra money for the privilege of having them fulfill the tasks the way you want them to.
Or just use the tools directly since they no longer underperform human actors. GPT-4 going toe to toe with the experts who set the benchmarks, way overperforming the crowdworkers https://www.artisana.ai/articles/gpt-4-outperforms-elite-cro... GPT-3.5 outperforming crowdworkers https://arxiv.org/abs/2303.15056 More money won't stop people from automating whatever they can. Especially when the automation provides such…
Don't think that'll work for training.
Humans go insane if we lack sufficiently varied sensory input. Dreaming doesn't really help with that, and I recall a recent paper that showed something similar for AI.
And even fine-tuning using AI causes the models to quickly degrade I've seen here previously. Perhaps we are already at peak of these LLMs now and the next gen won't come because all training data will be poisened with AI generate crap. That would be kinda funny to me tbh
I have a feeling that some types of data will become a bit like "the DNA wealth of the rainforest", valuable because it's so rich, different and rare. In particular small and/or dying human languages. Languages carry with them an immense amount of information, from all the precious human experiences that have shaped them.
Languages at fine granularity and narrative complexes at high granularity. Some people explicitly refer to stories as the most efficient compression mechanisms there is, by a large margin. There is a reason the Hutter prize is what it is. Which, in a supreme irony for the positivist types, rationally entails that folk tales and the Bible (gasp) are the most valuable data there is.