Live data from Hacker News

Anyone got a contact at OpenAI. They have a spider problem

mailman.nanog.org

201–210 of 400 posts

Re: Anyone got a contact at OpenAI. They have a spider problem

#201

Earlier quoted context omitted.

The end state of training on web text has always been an ouroboros - primarily because of adtech incentives to produce low quality content at scale to capture micro pennies. The irony of the whole thing is brutal.

Content you’re allowed and capable of scraping on the Internet is such a small amount of data, not sure why people are acting otherwise. Common crawl alone is only a few hundred TB, I have more content than that on a NAS sitting in my office that I built for a few grand (Granted I’m a bit of a data hoarder). The fears that we have “used all the data” are incredibly unfounded.

"Content you're allowed to scrape from the internet" is MUCH smaller than what LLMs have actually scraped, but they don't care about copyright.

> The fears that we have “used all the data” are incredibly unfounded.

The problem isn't whether we used all the real data or not, the problem is that it becomes increasingly difficult to distinguish real data from previous LLM outputs.

Re: Anyone got a contact at OpenAI. They have a spider problem

#202
post #104

Earlier quoted context omitted.

The end state of training on web text has always been an ouroboros - primarily because of adtech incentives to produce low quality content at scale to capture micro pennies. The irony of the whole thing is brutal.

> The end state of training on web text has always been an ouroboros And when other mediums have been saturated with AI? Books, music, radio, podcasts, movies -- what then? Do we need a (curated?) unadulterated stockpile of human content to avoid the enshittification of everything?

https://en.wikipedia.org/wiki/Low_background_steel but for web content.

Re: Anyone got a contact at OpenAI. They have a spider problem

#203
post #71

Eventually, OpenAI (and friends) are going to be training their models on almost exclusively AI generated content, which is more often than not slightly incorrect when it comes to Q&A, and the quality of AI responses trained on that content will quickly deteriorate. Right now, most internet content is written by humans. But in 5 years? Not so much. I think this is one of the big problems that the AI space needs to so…

It is already solved. Look at how Microsoft trained Phi - they used existing models to generate synthetic data from textbooks. That allowed them to create a new dataset grounded in “fact” at a far higher quality than common crawl or others.

It looks less like an ouroboros and more like a bootstrapping problem.

Re: Anyone got a contact at OpenAI. They have a spider problem

#204

Earlier quoted context omitted.

I've wondered if one could train a LLM on a closed set of curated knowledge. Then include training data that models the behaviour of not knowing. To the point that it could generalize to being able to represent its own not knowing. Because expecting a behaviour, like knowing you don't know, that isn't represented in the training set is silly. Kids make stuff up at first, then we correct them - so they have a way to l…

> I've wondered if one could train a LLM on a closed set of curated knowledge. Then include training data that models the behaviour of not knowing. To the point that it could generalize to being able to represent its own not knowing. The problem is that curating data is slow and expensive and downloading the entire web is fast and cheap. See also https://en.wikipedia.org/wiki/Cyc

Agreed. Using LLM to generate or curate training sets for other generations seems like a cool approach.

Maybe if you trained a small base model to know it doesn't know in general and THEN trained it on the entire web with embedded not-knowing preserving training examples, it would work?

Re: Anyone got a contact at OpenAI. They have a spider problem

#205
post #71

Eventually, OpenAI (and friends) are going to be training their models on almost exclusively AI generated content, which is more often than not slightly incorrect when it comes to Q&A, and the quality of AI responses trained on that content will quickly deteriorate. Right now, most internet content is written by humans. But in 5 years? Not so much. I think this is one of the big problems that the AI space needs to so…

I really really hope that five years from now we are not still using AI systems that behave the way today's do, based on probabilistic amalgamations of the whole of the internet. I hope we have designed systems that can reason about what they are learning and build reasonable mental models about what information is valuable and what can be discarded.

Re: Anyone got a contact at OpenAI. They have a spider problem

#206

Earlier quoted context omitted.

> I've wondered if one could train a LLM on a closed set of curated knowledge. Then include training data that models the behaviour of not knowing. To the point that it could generalize to being able to represent its own not knowing. The problem is that curating data is slow and expensive and downloading the entire web is fast and cheap. See also https://en.wikipedia.org/wiki/Cyc

Agreed. Using LLM to generate or curate training sets for other generations seems like a cool approach. Maybe if you trained a small base model to know it doesn't know in general and THEN trained it on the entire web with embedded not-knowing preserving training examples, it would work?

Reminded of this approach where a tiny model was trained children's stories generated by a larger model:

https://www.quantamagazine.org/tiny-language-models-thrive-w...

Re: Anyone got a contact at OpenAI. They have a spider problem

#207

Earlier quoted context omitted.

Yup, that's my impression as well. He's just nice to let OpenAI they have a problem. Usually this should be rewarded with a nice "hey, u guys have a bug" bounty because not long time ago some VP from OpenAI was lamenting that training their AI is, and it's his direct quote, "eye watering" cost (the order was millions of $$ per second).

I would be a little sceptical about that figure. 3 million dollars per second is around the world GDP. I get it, AI training is expensive, but I don't believe it's that expensive

Sam Altman has said in a few interviews that it was around $100 million for GPT-3, and higher for GPT-4.

But yes, this is a one-time cost, and far lower than the "millions of dollars per second" in GP comment.

https://fortune.com/2024/04/04/ai-training-costs-how-much-is...

https://www.wired.com/story/openai-ceo-sam-altman-the-age-of...

Re: Anyone got a contact at OpenAI. They have a spider problem

#208
post #86

Earlier quoted context omitted.

I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs. I mean if someone posts a question on an internet forum that I don't know the answer to, I'm certainly not going to post "I don't know" since that wouldn't be useful. In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence…

Reminds me of a joke Three logicians walk into a bar. The bartender says "what'll it be, three beers?" The first logician says "I don't know". The second logician says "I don't know". The third logician says "Yes".

And the human bartender passes the check to the third logician.

Re: Anyone got a contact at OpenAI. They have a spider problem

#209
post #71

Eventually, OpenAI (and friends) are going to be training their models on almost exclusively AI generated content, which is more often than not slightly incorrect when it comes to Q&A, and the quality of AI responses trained on that content will quickly deteriorate. Right now, most internet content is written by humans. But in 5 years? Not so much. I think this is one of the big problems that the AI space needs to so…

It is already solved. Look at how Microsoft trained Phi - they used existing models to generate synthetic data from textbooks. That allowed them to create a new dataset grounded in “fact” at a far higher quality than common crawl or others. It looks less like an ouroboros and more like a bootstrapping problem.

AI training on AI-generated content is a future problem. Using textbooks is a good idea, until our textbooks are being written by AI.

This problem can't really be avoided once we begin using AI to write, understand, explain, and disseminate information for us. It'll be writing more than blogs and SEO pages.

How long before we start readily using AI to write academic journals and scientific papers? It's really only a matter of time, if it's not already happening.

Re: Anyone got a contact at OpenAI. They have a spider problem

#210
post #86

Earlier quoted context omitted.

I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs. I mean if someone posts a question on an internet forum that I don't know the answer to, I'm certainly not going to post "I don't know" since that wouldn't be useful. In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence…

A lot of LLM hallucination is because of the internal conflict between alignment for helpfulness and lack of a clear answer. It's much like when someone gets out of their depth in a conversation and dissembles their way through to try and maintain their illusion of competence. In these cases, if you give the LLM explicit permission to tell you that it doesn't know in cases where it's not sure, that will significantly…

It also has a problem with quantity, so it gets confused by things like the cube root of 750l that it maintains for a long time is around 9m. It even suggests that 1l is equal to 1m³.
Post reply on HN