Live data from Hacker News

On the dangers of stochastic parrots: Can language models be too big? (2021)

dl.acm.org

41–50 of 111 posts

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#41
post #31
post #23

Earlier quoted context omitted.

> These machine dont think, and they dont understand But they do solve many tasks correctly, even problems with multiple steps and new tasks for which they got no specific training. They can combine skills in new ways on demand. Call it what you want.

They don't. Solve tasks, I mean. There's not a single task you can throw at them and rely on the answer. Could they solve tasks? Potentially. But how would we ever know that we could trust them? With humans we not only have millennia of collective experience when it comes to tasks, judging the result, and finding bullshitters. Also, we can retrain a human on the spot and be confident they won't immediately forget som…

> Also, we can retrain a human on the spot and be confident they won't immediately forget something important over that retraining.

I don’t have millennia, but my more than 3 decades of experience interacting with human beings tell me this is not nearly as reliable as you make it seem.

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#42
post #36

This paper is the product of a failed model of AI safety, in which dedicated safety advocates act as a public ombudsman with an adversarial relationship with their employer. It's baffling to me why anyone thought that would be sustainable. Compare this to something like RLHF[0] which has acheived far more for aligning models toward being polite and non-evil. (This is the technique that helps ChatGPT decline to answer…

> Compare this to something like RLHF[0] which has acheived far more for aligning models toward being polite and non-evil. (This is the technique that helps ChatGPT decline to answer questions like "how to make a bomb?") I recently saw a screenshot of someone doing trolley problems with people of all races & ages with ChatGPT and noting differences. That makes me not quite as confident about alignment as you are.

I am curious to see that trolley problem screenshot. I saw another screenshot where ChatGPT was coaxed into justifying gender pay differences by prompting it to generate hypothetical CSV or JSON data.

Basically you have to convince modern models to say bad stuff using clever hacks (compared to GPT-2 or even early GPT-3 where it would just spout straight-up hatred with the lightest touch).

That's very good progress and I'm sure there is more to come.

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#43
post #36

Earlier quoted context omitted.

> Compare this to something like RLHF[0] which has acheived far more for aligning models toward being polite and non-evil. (This is the technique that helps ChatGPT decline to answer questions like "how to make a bomb?") I recently saw a screenshot of someone doing trolley problems with people of all races & ages with ChatGPT and noting differences. That makes me not quite as confident about alignment as you are.

I am curious to see that trolley problem screenshot. I saw another screenshot where ChatGPT was coaxed into justifying gender pay differences by prompting it to generate hypothetical CSV or JSON data. Basically you have to convince modern models to say bad stuff using clever hacks (compared to GPT-2 or even early GPT-3 where it would just spout straight-up hatred with the lightest touch). That's very good progress an…

> I saw another screenshot where ChatGPT was coaxed into justifying gender pay differences by prompting it to generate hypothetical CSV or JSON data.

I remember seeing that on Twitter. My impression was author instructed the AI to discriminate by gender.

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#44
Pure speculation ahead-

The other day on Hacker News, there was that article about how scientists could not tell GPT-generated paper abstracts from real ones.

Which makes me think- abstracts for scientific papers are high-effort. The corpus of scientific abstracts would understandably have a low count of "garbage" compared to, say, Twitter posts or random blogs.

That's not to say that all scientific abstracts are amazing, just that their goal is to sound intelligent and convincing, while probably 60% of the junk fed into GPT is simply clickbait and junk content padded to fit some publisher's SEO requirements.

In other words, ask GPT to generate an abstract, and I would expect it to be quite good.

Ask it to generate a 5-paragraph essay about Huckleberry Finn, and I would expect it to be the same quality as the corpus- that is to say, high-school English students.

So now that we know these models can learn many one-shot tasks, perhaps some cleanup of the training data is required to advance. Imagine GPT trained ONLY on the library of congress, without the shitty travel blogs or 4chan rants.

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#45
post #6

This paper is embarrassingly bad. It's really just an opinion piece where the authors rant about why they don't like large language models. There is no falsifiable hypothesis to be found in it. I think this paper will age very poorly, as LLMs continue to improve and our ability to guide them (such as with RLHF) improves.

Why would there be a falsifiable hypothesis in it? Do you think that's a criterion for something being a scientific paper or something? If it ain't Popper, it ain't proper?

LLMs dramatically lower the bar for generating semi-plausible bullshit and it's highly likely that this will cause problems in the not-so-distant future. This is already happening. Ask any teacher anywhere. Students are cheating like crazy, letting chatGPT write their essays and answer their assignments without actually engaging with the material they're supposed to grok. News sites are pumping out LLM-generated articles and the ease of doing so means they have an edge over those who demand scrutiny and expertise in their reporting—it's not unlikely that we're going to be drowning in this type of content.

LLMs aren't perfect. RLHF is far from perfect. Language models will keep making subtle and not-so-subtle mistakes and dealing with this aspect of them is going to be a real challenge.

Personally, I think everyone should learn how to use this new technology. Adapting to it is the only thing that makes sense. The paper in question raised valid concerns about the nature of (current) LLMs and I see no reason why it should age poorly.

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#46
post #7

The problems with LLM are numerous but whats really wild to me is that even as they get better at fairly trivial tasks the advertising gets more and more out of hand. These machine dont think, and they dont understand, but people like the CEO of OpenAI allude to them doing just that, obviously so the hype can make them money.

> These machine dont think

And submarines don't swim.

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#47

This paper is the product of a failed model of AI safety, in which dedicated safety advocates act as a public ombudsman with an adversarial relationship with their employer. It's baffling to me why anyone thought that would be sustainable. Compare this to something like RLHF[0] which has acheived far more for aligning models toward being polite and non-evil. (This is the technique that helps ChatGPT decline to answer…

Isn't RLHF trivially easy to defeat (as it stands now)?

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#48
post #32

Earlier quoted context omitted.

This is going to happen alot over the next few years. One can fine tune GPT-2 medium on an RTX2070. Training GPT-2 medium from scratch can be done for $162 on vast.ai. The newer H100/Trainium/Tensorcore chips will bring the price down even further. I suspect if one wanted to fully replicate ChatGPT from scratch it would take ~1-2 million including label acquisition. You probably only require ~200-500k in compute. The…

These things have reached the tipping point where they provide significant utility to a significant portion of the computer scientists working on making these things. Could be that the coming iterations of these new tools will make it increasingly easy to write the code for the next iterations of these tools. I wonder if this is the first rumblings of the singularity.

I can imagine a world where there are an infinity of “local maximums” that stop a system from reaching a singular feedback loop… imagine if our current tools help write the next generation, so on, so on, until it gets stuck in some local optimization somewhere. Getting stuck seems more likely than not getting stuck, right?

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#49

Pure speculation ahead- The other day on Hacker News, there was that article about how scientists could not tell GPT-generated paper abstracts from real ones. Which makes me think- abstracts for scientific papers are high-effort. The corpus of scientific abstracts would understandably have a low count of "garbage" compared to, say, Twitter posts or random blogs. That's not to say that all scientific abstracts are ama…

The science is in the reproduction of the methodology, not in the abstract… in fact, a lot of garbage publications with catchy abstracts built on a shaky foundation sounds like one of the issues that plagues contemporary science. That people would stop finding abstracts useful seems a good thing!

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#50
I'm still midway through the paper, but I gotta say, I'm a little surprised at the contrast between the contents of the paper and how people have described it on HN. I don't agree with everything that is said, but there are some interesting points made about the data used to train the models, such as it capturing bias (I would certainly question the methodology of using reddit as a large source of training data), and that bias being amplified by filtering algorithms that produce the even larger datasets used for modern LLMs. The section about environmental impact might not hit home for everyone, but it is valid to raise issues around the compute usage involved in training these models. First, because it limits this training to companies who can spend millions of dollars on compute, and second because if we want to scale up models, efficiency is probably a top goal.

What really confuses me here is how this paper is somehow outside the realm of valid academic discourse. Yes, it is steeped in activist, social justice language. Yes, it has a different perspective than most CS papers. But is that wrong? Is that enough of a sin to warrant such a response that this paper has received? I'll need to finish the paper to fully judge, but I'm leaning towards no, it is not enough of a sin.

Post reply on HN