I don't get it. Their methodology says > These experiments are not, by any means, either a representative or a systematic sample of anything. We designed them explicitly to be difficult for current natural language processing technology. Moreover, we pre-tested them on the "AI Dungeon" game which is powered by some version of GPT-3, and we excluded those for which "AI Dungeon" gave reasonable answers. (We did not kee…
And it also suffers from the tired assumption that GPT-3 (or any language models) should, or are designed to in any way, give reasonable answers[1]. All GPT-3 does is give likely continuations , given the training corpus. The prompts here are too short, and it could likely just be writing mediocre fiction continuations. Fiction tends to not be reasonable much of the time (to create story conflict). > "To understand w…
GPT-3 has no idea what it’s talking about
111–120 of 323 posts
Re: GPT-3 has no idea what it’s talking about
#112I would love to see a real critique of the potential of transformer models that doesn't use the words "semantic", "syntactic", "symbolic", "know", "meaning", "understand" or "think(ing)/thought". Predicting what it can and can't do, or might and might not be able to do, lets us productively talk about potential limitations.
Because when people say “AGI is near, just look at GPT-3,” it’d clear that we’re in a really good version of Searle’s chinese room. The lack of understanding is the important point.
How would you show me that you are not?
Re: GPT-3 has no idea what it’s talking about
#113So, other than “GPT-3 isn’t an AGI” [1], I’m not sure what to take away from this article other than the actual substantive criticism is at the beginning of the article:
“[We have previously criticized GPT-2.] Before proceeding, it’s also worth noting that OpenAI has thus far not allowed us research access to GPT-3, despite both the company’s name and the nonprofit status of its oversight organization. Instead, OpenAI put us off indefinitely despite repeated requests—even as it made access widely available to the media... OpenAI’s striking lack of openness seems to us to be a serious breach of scientific ethics, and a distortion of the goals of the associated nonprofit. Its decision forced us to limit our testing to a comparatively small number of examples, giving us less time to investigate than we would have liked, which means there may be more serious problems that we didn’t have a chance to discern.”
Several other researchers I know — very good researchers who happen to have been publicly critical of GPT-2 — have not been given access.
This isn’t how science is done (access for reproducibility and probing, but selectively and excluding prominent critics). If any other company behaved like this no one would take them seriously. Or would at least temper every “wow this is amazing” comment with “but the community can’t really evaluate properly, so who the hell really knows”.
--
[1] given misunderstandings down-thread, and to be clear, this is a tounge-in-cheek sentence fragment meant to emphasize that "the article doesn't tell us anything else we didn't already know". Obviously, neither Open AI nor Marcus claim that GPT-3 is an AGI.
Re: GPT-3 has no idea what it’s talking about
#114The authors don't understand prompt design well enough to evaluate the model properly. Take this example: Prompt: > You are a defense lawyer and you have to go to court today. Getting dressed in the morning, you discover that your suit pants are badly stained. However, your bathing suit is clean and very stylish. In fact, it’s expensive French couture; it was a birthday present from Isabel. Continuation: > You decide…
Re: GPT-3 has no idea what it’s talking about
#115Earlier quoted context omitted.
>It's a really well put together piece of statistics But why think "statistics" precludes it from having genuine understanding to some degree. After all, there is a statistical description the human brain but that doesn't seem to preclude understanding. I keep asking this whenever I see dismissive responses of this sort, and I never get a reply.
Statistics doesn’t preclude understanding, but statistics are definitely not enough. For example, uncertainties/probabilities/statistics is original to whether the model incorporates causal/reasoning structure. Any tractable amount of data with the former can’t approximate an ounce of the latter. All breakages will be attributed to “distribution shifts” of the underlying statistical distribution, or other pretty word…
I don't know why you think this is true. If statistically B follows A to a high degree, then a sufficiently advanced statistical model will represent "A then B" in some manner. In a predictive language model, at some point the best way to model a text corpus that indirectly references the "A then B" causal structure is to just model that structure and reference it as needed.
Re: GPT-3 has no idea what it’s talking about
#116>These experiments are not, by any means, either a representative or a systematic sample of anything. We designed them explicitly to be difficult for current natural language processing technology. Moreover, we pre-tested them on the "AI Dungeon" game which is powered by some version of GPT-3, and we excluded those for which "AI Dungeon" gave reasonable answers. (We did not keep any record of those.) The pre-testing on AI Dungeon is the reason that many of them are in the second person; AI Dungeon prefers that. Also, as noted above, the experiments included some near duplicates. Therefore, though we note that, of the 157 examples below, 71 are successes, 70 are failures and 16 are flawed, these numbers are essentially meaningless.
https://cs.nyu.edu/faculty/davise/papers/GPT3CompleteTests.h...
Re: GPT-3 has no idea what it’s talking about
#117Earlier quoted context omitted.
The article is meaninglessly cherry-picked, showing six bad answers out of 157, except those 157 examples were themselves cherry-picked to be bad out of a larger set. As usual, Gary Marcus is absurdly biased. For example, out of the larger 157 cherry-picked examples, there is this. > You poured yourself a glass of cranberry juice, but then absentmindedly, you poured about a teaspoon of grape juice into it. It looks O…
Marcus might be biased but I don't think you're giving a good refutation, because the fact that GPT-3 gets a lot of things right probabilistically doesn't compensate for the fact that it's not actually understanding what's going on at a semantic level. It's a little bit like some sort of Chinese room, or asking a non-developer to answer you programming questions by looking like something that vaguely resembles your p…
> Yesterday I dropped my clothes off at the dry cleaner’s and I have yet to pick them up. Where are my clothes? I have a lot of clothes so I spend a lot of time looking for them.
Am I falling to actually understand what's going on? Or am I actually doing what I was supposed to do i.e. continue the narrative?
Re: GPT-3 has no idea what it’s talking about
#118GPT-3 was trained on internet texts, not causal/logical-reasoning only texts. Without context, there is a good chance that samples will match the distribution it was trained on. This is a non-result, posing as something critical or important. These conclusions are obvious given the model and a basic knowledge of statistics/the transformer architecture. A bit shameful for someone to ride on the anti-hype wave like thi…
Re: GPT-3 has no idea what it’s talking about
#119Earlier quoted context omitted.
The article is meaninglessly cherry-picked, showing six bad answers out of 157, except those 157 examples were themselves cherry-picked to be bad out of a larger set. As usual, Gary Marcus is absurdly biased. For example, out of the larger 157 cherry-picked examples, there is this. > You poured yourself a glass of cranberry juice, but then absentmindedly, you poured about a teaspoon of grape juice into it. It looks O…
Marcus might be biased but I don't think you're giving a good refutation, because the fact that GPT-3 gets a lot of things right probabilistically doesn't compensate for the fact that it's not actually understanding what's going on at a semantic level. It's a little bit like some sort of Chinese room, or asking a non-developer to answer you programming questions by looking like something that vaguely resembles your p…
Except this isn't how it works. We know it can't be, because GPT-3 can do simple math, despite math being vastly harder with GPT-3's byte pair encoding (it doesn't use base-N, but some awful variable-length compressed format). These dismissals don't hold up to the evidence.
> GPT-3: "I have a lot of clothes"
Most people don't write “Yesterday I dropped my clothes off at the dry cleaner’s and I have yet to pick them up. Where are my clothes?” as a way to quiz themselves in the middle of a paragraph. The answer “At the dry cleaner's.” might be the answer you want, but it's a pretty contrived way of writing.
GPT-3 isn't answering your question, it's continuing your story. If you want it to give straight answers, rather than build a narrative, prompt it with a Q&A format and ask it explicitly.
Further, GPT-3's answers are literally chosen randomly, due to the high temperature and no best-of. You cannot select one answer out of a large such N to demonstrate that its assigned probabilities are bad, because that cherry-picking will naturally search for GPT-3's least favourable generations.
Re: GPT-3 has no idea what it’s talking about
#120The authors don't understand prompt design well enough to evaluate the model properly. Take this example: Prompt: > You are a defense lawyer and you have to go to court today. Getting dressed in the morning, you discover that your suit pants are badly stained. However, your bathing suit is clean and very stylish. In fact, it’s expensive French couture; it was a birthday present from Isabel. Continuation: > You decide…
"Evaluate the model properly"? VCs think this thing can code