Live data from Hacker News

On the dangers of stochastic parrots: Can language models be too big? (2021)

dl.acm.org

1–10 of 111 posts

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#6
This paper is embarrassingly bad. It's really just an opinion piece where the authors rant about why they don't like large language models.

There is no falsifiable hypothesis to be found in it.

I think this paper will age very poorly, as LLMs continue to improve and our ability to guide them (such as with RLHF) improves.

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#7
The problems with LLM are numerous but whats really wild to me is that even as they get better at fairly trivial tasks the advertising gets more and more out of hand. These machine dont think, and they dont understand, but people like the CEO of OpenAI allude to them doing just that, obviously so the hype can make them money.

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#8
post #7

The problems with LLM are numerous but whats really wild to me is that even as they get better at fairly trivial tasks the advertising gets more and more out of hand. These machine dont think, and they dont understand, but people like the CEO of OpenAI allude to them doing just that, obviously so the hype can make them money.

Could be the sign of the next AI winter.

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#9
This was mostly political guff about environmentalism and bias, but one thing I didn't know was that apparently larger models make it easier to extract training data.

> Finally, we note that there are risks associated with the fact that LMs with extremely large numbers of parameters model their training data very closely and can be prompted to output specific information from that training data. For example, [28] demonstrate a methodology for extracting personally identifiable information (PII) from an LM and find that larger LMs are more susceptible to this style of attack than smaller ones. Building training data out of publicly available documents doesn’t fully mitigate this risk: just because the PII was already available in the open on the Internet doesn’t mean there isn’t additional harm in collecting it and providing another avenue to its discovery. This type of risk differs from those noted above because it doesn’t hinge on seeming coherence of synthetic text, but the possibility of a sufficiently motivated user gaining access to training data via the LM. In a similar vein, users might query LMs for ‘dangerous knowledge’ (e.g. tax avoidance advice), knowing that what they were getting was synthetic and therefore not credible but nonetheless representing clues to what is in the training data in order to refine their own search queries

Shame they only gave that one graf. I'd like to know more about this. Again, miss me with the political garbage about "dangerous knowledge", the most concerning thing is the PII leakage as far as I can tell.

Re: On the dangers of stochastic parrots: Can language models be too big? (2021)

#10
post #6

This paper is embarrassingly bad. It's really just an opinion piece where the authors rant about why they don't like large language models. There is no falsifiable hypothesis to be found in it. I think this paper will age very poorly, as LLMs continue to improve and our ability to guide them (such as with RLHF) improves.

[deleted]
Post reply on HN