On the dangers of stochastic parrots: Can language models be too big? (2021)
1–10 of 111 posts
Re: On the dangers of stochastic parrots: Can language models be too big? (2021)
#2Re: On the dangers of stochastic parrots: Can language models be too big? (2021)
#3Re: On the dangers of stochastic parrots: Can language models be too big? (2021)
#4I believe this is the papers that got timnit and mmitchel fired from google, followed by a protracted media/legal campaign against google and vice versa.
Re: On the dangers of stochastic parrots: Can language models be too big? (2021)
#5Re: On the dangers of stochastic parrots: Can language models be too big? (2021)
#6There is no falsifiable hypothesis to be found in it.
I think this paper will age very poorly, as LLMs continue to improve and our ability to guide them (such as with RLHF) improves.
Re: On the dangers of stochastic parrots: Can language models be too big? (2021)
#7Re: On the dangers of stochastic parrots: Can language models be too big? (2021)
#8The problems with LLM are numerous but whats really wild to me is that even as they get better at fairly trivial tasks the advertising gets more and more out of hand. These machine dont think, and they dont understand, but people like the CEO of OpenAI allude to them doing just that, obviously so the hype can make them money.
Re: On the dangers of stochastic parrots: Can language models be too big? (2021)
#9> Finally, we note that there are risks associated with the fact that LMs with extremely large numbers of parameters model their training data very closely and can be prompted to output specific information from that training data. For example, [28] demonstrate a methodology for extracting personally identifiable information (PII) from an LM and find that larger LMs are more susceptible to this style of attack than smaller ones. Building training data out of publicly available documents doesn’t fully mitigate this risk: just because the PII was already available in the open on the Internet doesn’t mean there isn’t additional harm in collecting it and providing another avenue to its discovery. This type of risk differs from those noted above because it doesn’t hinge on seeming coherence of synthetic text, but the possibility of a sufficiently motivated user gaining access to training data via the LM. In a similar vein, users might query LMs for ‘dangerous knowledge’ (e.g. tax avoidance advice), knowing that what they were getting was synthetic and therefore not credible but nonetheless representing clues to what is in the training data in order to refine their own search queries
Shame they only gave that one graf. I'd like to know more about this. Again, miss me with the political garbage about "dangerous knowledge", the most concerning thing is the PII leakage as far as I can tell.
Re: On the dangers of stochastic parrots: Can language models be too big? (2021)
#10This paper is embarrassingly bad. It's really just an opinion piece where the authors rant about why they don't like large language models. There is no falsifiable hypothesis to be found in it. I think this paper will age very poorly, as LLMs continue to improve and our ability to guide them (such as with RLHF) improves.