The AI bullshit singularity
81–90 of 187 posts
Re: The AI bullshit singularity
#82LLMs are increasingly trained on synthetic data. A LLM can learn logic, and ultimately the rules governing physics and everything else, without human input. The first versions being trained on "human knowledge" seems kind of like a proof-of-concept. Future iterations will be much, much smarter.
Re: The AI bullshit singularity
#83I always found the idea of infinitely self improving AI to be suspect. Let’s say we have a super smart AI with intelligence 1, and it uses all that to improve itself by 0.5. Then that new 1.5 uses itself to improve by 0.25. Then 0.125, etc etc. obviously it’s always increasing, but it’s not going to have the runaway effect people think.
Saying that models will get attracted to bullshit local maximum is similar fallacy to saying that wikipedia will be full of rubbish when it was created. Forces are set up in a way that creates improvements that accumulate, humans don't represent any ceiling and unlike humans models have near zero replication cost, especially time wise.
Re: The AI bullshit singularity
#84Re: The AI bullshit singularity
#85Earlier quoted context omitted.
Evolution has a clearly defined fitness function (number of offspring that are produced). What is the fitness function for generated art?
Number of works that are released to the internet! At the end of the day, there is a human who has an idea in mind and is using an LLM to realize it. They are tweaking their prompt until they achieve their vision, then publishing only those successes. The resulting art can then be used with the prompt as training data.
Likewise SEO incentivizes mass production of low quality LLM text, because they "quality" they are optimizing for is impressions, not actual quality.
Re: The AI bullshit singularity
#86And everybody's asking what AI is going to be and can do for them, all the while freely working for it. Don't ask what AI can do for you, just do for AI what is asked of you.
Re: The AI bullshit singularity
#87Succinctly stated and something that resonates strongly with me. In the last internet revolution (web search), results started high quality because the inputs were high quality - bloggers and others just wanted to document and share knowledge. But over time, many interests (largely commercial) figured out how to game the system with SEO, and quality of search results has decreased as search's incentive structure led…
I think it’s clear that LLMs cannot be the end state of this technology, and we will need systems that can reason and develop hypotheses and test them internally. These systems may benefit from more curated datasets (such as those collected before the bullshit wave began) along with real world interaction data from YouTube and robotics. Such systems could eventually be used to rank web pages for their bullshit level,…
Since all deployed models produce some kind of effect and feedback from the world there is an opportunity there to collect data targeted on the current level of the model, the most useful kind of data. That's why I think AI will be ok even with the proliferation of bots online. It's not 100% pure synthetic data in a loop, it is a agent-environment loop.
tl;dr Models learn better from their own experiences, not ours.
Re: The AI bullshit singularity
#88Succinctly stated and something that resonates strongly with me. In the last internet revolution (web search), results started high quality because the inputs were high quality - bloggers and others just wanted to document and share knowledge. But over time, many interests (largely commercial) figured out how to game the system with SEO, and quality of search results has decreased as search's incentive structure led…
We may enter into a highly ironic regime where ai companies subsidize writers and artists at scale to produce better data.
Re: The AI bullshit singularity
#89Re: The AI bullshit singularity
#90I broadly agree with that. Repeated training with self-generated data is the technological equivalent of incest and can lead to nothing good.
Humans train on self-generated data all the time, but compared to today's LLMs, humans have superior reasoning ability, which enables them to separate good from bad data. Humans can confirm many things for themselves if need be and our social hierarchies crystallize authority figures which perform curation for others. On top of that humans have a lifetime of experience with the real world, while LLMs rely only on rep…
I think people are substantially worse at figuring out what is “good” data, and I think a huge number of people are/are starting to nominate LLMs as those “authority figures”. And, given that, I think there is a low ceiling for the creators of these tools to clear to make them appealing to users.