> Why is this significant? At the core the model is still doing language modeling, right? learning to predict the next word, based on text alone? Sure, but here the human annotators inject some level of grounding to the text. Some symbols ("summarize", "translate", "formal") are used in a consistent way together with the concept/task they denote. And they always appear in the beginning of the text. This make these symbols (or the "instructions") in some loose sense external to the rest of the data, making the act of producing a summary grounded to the human concept of "summary". Or in other words, this helps the model learn the communicative intent of the a user who asks for a "summary" in its "instruction". An objection here would be that such cases likely naturally occur already in large text collections, and the model already learned from them, so what is new here? I argue that it might be much easier to learn from direct instructions like these than it is to learn from non-instruction data (think of a direct statement like "this is a dog" vs needing to infer from over-hearing people talk about dogs). And that by shifting the distribution of the training data towards these annotated cases, substantially alter how the model acts, and the amount of "grounding" it has. And that maybe with explicit instructions data, we can use much less training text compared to what was needed without them. (I promised you hand waving didn't I?)
Some Remarks on Large Language Models
11–20 of 87 posts
Re: Some Remarks on Large Language Models
#12The dismissal of biases and stereotypes is exactly why AI research needs more people who are part of the minority. Yoav can dismiss this because it just doesn't affect him much. It's easy to say "Oh well, humans are biased too" when the biases of these machines don't: misgender you, mistranslate text that relates to you, have negative affect toward you, are more likely to write violent stories related to you, have lo…
I mean he is presumably Jewish and lives in Israel, so I would guess he knows quite a bit about being a minority and experiencing bias.
Re: Some Remarks on Large Language Models
#13Re: Some Remarks on Large Language Models
#14The dismissal of biases and stereotypes is exactly why AI research needs more people who are part of the minority. Yoav can dismiss this because it just doesn't affect him much. It's easy to say "Oh well, humans are biased too" when the biases of these machines don't: misgender you, mistranslate text that relates to you, have negative affect toward you, are more likely to write violent stories related to you, have lo…
I mean he is presumably Jewish and lives in Israel, so I would guess he knows quite a bit about being a minority and experiencing bias.
Re: Some Remarks on Large Language Models
#15I found the "grounding" explanation provided by human feedback very insightful: > Why is this significant? At the core the model is still doing language modeling, right? learning to predict the next word, based on text alone? Sure, but here the human annotators inject some level of grounding to the text. Some symbols ("summarize", "translate", "formal") are used in a consistent way together with the concept/task they…
Re: Some Remarks on Large Language Models
#16The dismissal of biases and stereotypes is exactly why AI research needs more people who are part of the minority. Yoav can dismiss this because it just doesn't affect him much. It's easy to say "Oh well, humans are biased too" when the biases of these machines don't: misgender you, mistranslate text that relates to you, have negative affect toward you, are more likely to write violent stories related to you, have lo…
" The models encode many biases and stereotypes. Well, sure they do. They model observed human's language, and we humans are terrible beings, we are biased and are constantly stereotyping. This means we need to be careful when applying these models to real-world tasks, but it doesn't make them less valid, useful or interesting from a scientiic perspective." Not sure how this can be seen as dismissive. >Yoav can dismi…
Re: Some Remarks on Large Language Models
#17Sometimes I read text like this and really enjoy the deep insights and arguments once I filter out the emotion, attitude, or tone. And I wonder if the core of what they're trying to communicate would be better or more efficiently received if the text was more neutral or positive. E.g. you can be 'bearish' on something and point out 'limitations', or you can say 'this is where I think we are' and 'this is how I think…
Re: Some Remarks on Large Language Models
#18Earlier quoted context omitted.
I mean he is presumably Jewish and lives in Israel, so I would guess he knows quite a bit about being a minority and experiencing bias.
That's true, he probably sees a lot of anti-Palestinian bias on a regular basis.
Re: Some Remarks on Large Language Models
#19> Also, let's put things in perspective: yes, it is enviromentally costly, but we aren't training that many of them, and the total cost is miniscule compared to all the other energy consumptions we humans do.
Part of the reason LLMs aren't that big in the grand scheme of things is because they haven't been good enough and businesses haven't started to really adopt them. That will change, but the costs will be high because they're also extremely expensive to run. I think the author is focusing on the training costs for now, but that will likely get dwarfed by operational costs. What then? Waving one's arms and saying it'll just "get cheaper over time" isn't an acceptable answer because it's hard work and we don't really know how cheap we can get right now. It must be a focus if we actually care about widespread adoption and environmental impact.
Re: Some Remarks on Large Language Models
#20Earlier quoted context omitted.
" The models encode many biases and stereotypes. Well, sure they do. They model observed human's language, and we humans are terrible beings, we are biased and are constantly stereotyping. This means we need to be careful when applying these models to real-world tasks, but it doesn't make them less valid, useful or interesting from a scientiic perspective." Not sure how this can be seen as dismissive. >Yoav can dismi…
Or maybe he is blind to or unaffected by such biases either due to luck or wealth or other outliers. Especially as a Jewish person in Israel. There are always plenty of people in minority groups that feel (either correctly or incorrectly) that bias doesn't affect them. Take Clarence Thomas for example, or Candace Owens. Simply being a member of a minority group does not make your opinion correct. Thomas even said in…
This is what I really don’t like about the AI ethics critics (of the woke variety): it’s super easy to be dismissive, but it’s crazy hard to do anything that moves the world. If you move the world, some people will naturally be happy and others angry. Even creating a super “balanced” dataset will piss off those who want an imbalanced world!
No opinion is “correct” - they’re just opinions, including mine right now!