Live data from Hacker News

The AI bullshit singularity

successfulsoftware.net

81–90 of 187 posts

Re: The AI bullshit singularity

#82
post #42

LLMs are increasingly trained on synthetic data. A LLM can learn logic, and ultimately the rules governing physics and everything else, without human input. The first versions being trained on "human knowledge" seems kind of like a proof-of-concept. Future iterations will be much, much smarter.

[deleted]

Re: The AI bullshit singularity

#83

I always found the idea of infinitely self improving AI to be suspect. Let’s say we have a super smart AI with intelligence 1, and it uses all that to improve itself by 0.5. Then that new 1.5 uses itself to improve by 0.25. Then 0.125, etc etc. obviously it’s always increasing, but it’s not going to have the runaway effect people think.

There are many dimensions where improvements are happening - speed increase, size reduction, precision, context length, using external computation (function calling), using formal systems, hybrid setups, multi-modality etc. If you look at short history of what's happening - we're not seeing below 50% improvements over those relatively short periods of time. We had gpt1 just five and a half years ago. We now have open weight models orders of magnitude better. We know we're feeding models with tons of redundancy and low quality inputs, we know synthetic data can improve and lower training cost dramatically. We know we're not near anything optimal. We'll see orders of magnitude size reductions in coming years etc. Humans don't represent any kind of intelligence ceiling - it can be surpassed and if it can be surpassed and we know humans alone produce well above 50% improvements - it will get better and getting better.

Saying that models will get attracted to bullshit local maximum is similar fallacy to saying that wikipedia will be full of rubbish when it was created. Forces are set up in a way that creates improvements that accumulate, humans don't represent any ceiling and unlike humans models have near zero replication cost, especially time wise.

Re: The AI bullshit singularity

#84
Well, for now, anyway. I went to college in the 70’s. What strikes me is, once obviously wrong hallucinations are eliminated, LLM’s will do exactly what untalented liberal arts students did with their time, but with some useful stuff tacked on.

Re: The AI bullshit singularity

#85

Earlier quoted context omitted.

Evolution has a clearly defined fitness function (number of offspring that are produced). What is the fitness function for generated art?

Number of works that are released to the internet! At the end of the day, there is a human who has an idea in mind and is using an LLM to realize it. They are tweaking their prompt until they achieve their vision, then publishing only those successes. The resulting art can then be used with the prompt as training data.

It would be nice if every human at the wheel of an LLM or image generator were putting that much care into their craft, but that's just not what's happening. You just need to look at places like DeviantArt which are now overrun with users who joined 6 months ago and already have submissions in the thousands, it's simply not possible that they are putting any care into what they're spraying out. Much of the time they're posting numerous functionality identical images, probably generated from the same prompt with a different seed.

Likewise SEO incentivizes mass production of low quality LLM text, because they "quality" they are optimizing for is impressions, not actual quality.

Re: The AI bullshit singularity

#86
Of course it's not going to improve on itself. That's what you all are here for - happily interacting with LLM's. Doing the finetuning and improving it with yet another dose of humanity, albeit with more direct interaction this time.

And everybody's asking what AI is going to be and can do for them, all the while freely working for it. Don't ask what AI can do for you, just do for AI what is asked of you.

Re: The AI bullshit singularity

#87

Succinctly stated and something that resonates strongly with me. In the last internet revolution (web search), results started high quality because the inputs were high quality - bloggers and others just wanted to document and share knowledge. But over time, many interests (largely commercial) figured out how to game the system with SEO, and quality of search results has decreased as search's incentive structure led…

I think it’s clear that LLMs cannot be the end state of this technology, and we will need systems that can reason and develop hypotheses and test them internally. These systems may benefit from more curated datasets (such as those collected before the bullshit wave began) along with real world interaction data from YouTube and robotics. Such systems could eventually be used to rank web pages for their bullshit level,…

Not just real world data from videos. AI models need feedback from many sources: humans, code execution, web search, simulations, games, robotics, math verification, or from actual experiments in the real world. All of these are environments that can take the output of a model and do some processing and return feedback. The model can learn and search for solutions, creating its own training data as a RL agent.

Since all deployed models produce some kind of effect and feedback from the world there is an opportunity there to collect data targeted on the current level of the model, the most useful kind of data. That's why I think AI will be ok even with the proliferation of bots online. It's not 100% pure synthetic data in a loop, it is a agent-environment loop.

tl;dr Models learn better from their own experiences, not ours.

Re: The AI bullshit singularity

#88
post #68

Succinctly stated and something that resonates strongly with me. In the last internet revolution (web search), results started high quality because the inputs were high quality - bloggers and others just wanted to document and share knowledge. But over time, many interests (largely commercial) figured out how to game the system with SEO, and quality of search results has decreased as search's incentive structure led…

We may enter into a highly ironic regime where ai companies subsidize writers and artists at scale to produce better data.

I like this idea very much. Is this similar to human moderation of training data?

Re: The AI bullshit singularity

#89
LLMs are being trained on a smaller and smaller percentage of human prose. Right now it seems like code is the best source for the bulk of an LLM's diet, but it's also looking likely that synthetic math text will be even better. The structured reasoning of code and math seems to be what actually makes these big LLMs "smart." Once you've trained a smart LLM, it seems to take a relatively small amount of hand-curated human prose to fine tune it into talking like a human. Unfortunately this article feels like the wishful thinking of someone who is afraid of the changes LLMs are bringing and hasn't done much research.

Re: The AI bullshit singularity

#90
post #2

I broadly agree with that. Repeated training with self-generated data is the technological equivalent of incest and can lead to nothing good.

Humans train on self-generated data all the time, but compared to today's LLMs, humans have superior reasoning ability, which enables them to separate good from bad data. Humans can confirm many things for themselves if need be and our social hierarchies crystallize authority figures which perform curation for others. On top of that humans have a lifetime of experience with the real world, while LLMs rely only on rep…

This view of people seems… idealistic, given the number of people who uncritically look to Facebook or Reddit for information (and then happily parrot whatever memes they find) and the number of LLMs that do the same.

I think people are substantially worse at figuring out what is “good” data, and I think a huge number of people are/are starting to nominate LLMs as those “authority figures”. And, given that, I think there is a low ceiling for the creators of these tools to clear to make them appealing to users.

Post reply on HN