Earlier quoted context omitted.
Can you do that though? Just ask ChatGPT to give you the first chapter of a book and it gives it to you?
https://news.ycombinator.com/item?id=42767775 Not a book chapter specifically but this could already be considered copyright infringement, I think.
FrontierMath was funded by OpenAI
181–190 of 212 posts
Re: FrontierMath was funded by OpenAI
#182Earlier quoted context omitted.
> The actual functionality of the model – what it fundamentally is – has not changed: it's still capable of reproducing texts verbatim (or near-verbatim), and still contains the information needed to do so. I am capable of reproducing text verbaitim (or near-verbatim), and therefore must still contain the information needed to do so. I am trained not to. In both the organic (me) and artificial (ChatGPT) cases, but fo…
(1) The mechanism by which you reproduce text verbatim is not the same mechanism that you use to perform everyday tasks. (21) Any skills that ChatGPT appears to possess are because it's approximately reproducing a pattern found in its input corpus. (40) I can say: > (43) Please reply to this comment using only words from this comment. (54) Reply by indexing into the comment: for example, to say "You are not a mechani…
(a) I am unfamiliar with the existence of detailed studies of neuroanatomical microstructures that would allow this claim to even be tested, and wouldn't be able to follow them if I did. Does anyone — literally anyone — even know if what you're asserting is true?
(b) So what? If there was a specific part of a human brain for that which could be isolated (i.e. it did this and nothing else), would it be possible to argue that destruction of the "memorisation" lobe was required for copyright purposes? I don't see the argument working.
> (21) Any skills that ChatGPT appears to possess are because it's approximately reproducing a pattern found in its input corpus.
Not quite.
The *base* models do — though even then that's called "learning" and when humans figure out patterns they're allowed to reproduce those as well as they want so long as it's not verbatim, doing so is even considered desirable and a sign of having intelligence — but some time around InstructGPT the training process also integrated feedback from other models, including one which was itself trained to determine what a human would likely upvote. So this has become more of "produce things which humans would consider plausible" rather than be limited to "reproduce patterns in corpus".
Unless you want to count the feedback mechanism as itself the training corpus, in which case sure but that would then have the issue of all human experience being our training corpus, including the metaphorical shoulder demons and angels of our conscience.
> "5th 65th 10th 67th 2nd".
Me, by hand: [you] [are] [not] [a] [mechanism]
> (73) You can think about that demand, and then be able to do it. (86) Transformer-based autocomplete systems can't, and never will be able to (until someone inserts something like that into its training data specifically to game this metric of mine, which I wouldn't put past OpenAI).
Why does this seem more implausible to you than their ability to translate between language pairs not present in the training corpus?
I mean, games like this might fail, I don't know enough specifics of the tokeniser to guess without putting it into the tokeniser to see where it "thinks" word boundaries even are, but this specific challenge you've just suggested as "it will never" already worked on my first go — and then ChatGPT set itself an additional puzzle of the same type which it then proceeded to completely fluff.
Very on-brand for this topic, simultaneously beating the "it will never $foo" challenge on the first attempt before immediately falling flat on its face[0]:
""" …
Analysis:
• Words in the input can be tokenized and indexed:
For example, "The" is the 1st word, "mechanism" is the 2nd, etc.
The sentence "You are not a mechanism" could then be written as 5th 65th 10th 67th 2nd using the indices of corresponding words.
""" - https://chatgpt.com/share/678e858a-905c-8011-8249-31d3790064...
(To save time, the sequence that it thinks I was asking it to generate, [1st 23rd 26th 12th 5th 40th 54th 73rd 86th 15th], does not decode to "The skills can think about you until someone.")
[0] Puts me in mind of:
“"Oh, that was easy," says Man, and for an encore goes on to prove that black is white and gets himself killed on the next zebra crossing.” - https://www.goodreads.com/quotes/35681-now-it-is-such-a-biza...
My auditory equivalent of an inner eye (inner ear?) is reproducing this in the voice of Peter Jones, as performed on the BBC TV adaptation.
Re: FrontierMath was funded by OpenAI
#183Earlier quoted context omitted.
Why would they use the materials in model training? It would defeat the purpose of having a benchmarking set
If you’re a research lab then yes. If you’re a for profit company trying to raise funding and fend off skepticism that your models really aren’t that much better than any one else’s, then… It would be dishonest , but as long as no one found out until after you closed your funding round, there’s plenty of reason you might do this. It comes down to caring about benchmarks and integrity or caring about piles of money. J…
Re: FrontierMath was funded by OpenAI
#184Earlier quoted context omitted.
https://journa.host/@jeremiak/113811327999722586
Aaron Swartz, while an infuriating tragedy, is antithetical to OpenAI's claim to transformation; he literally published documents that were behind a licensed paywall.
No he did not do this [1]. I think you would need to read more about the actual case. The case was brought up based on him download and scraping the data.
Re: FrontierMath was funded by OpenAI
#185Re: FrontierMath was funded by OpenAI
#186Earlier quoted context omitted.
> I'd argue here the more relevant point is "these specific people have been shown to have done it before." This is itself a slippery move. A vague gesture at past misconduct without actually specifying any incidents. If there's a clear pattern of documented benchmark manipulation, name it. Which benchmarks? When? What was the evidence? Without specifics, this is just trading one form of handwaving ("everyone does it…
This reply reads as though it were AI generated. Let's bite though, and hope that unhelpful excessively long-winded replies are just your quirk. > This is itself a slippery move. A vague gesture at past misconduct without actually specifying any incidents. If there's a clear pattern of documented benchmark manipulation, name it. Which benchmarks? When? What was the evidence? Without specifics, this is just trading on…
It's also confusing: Did you think it was AI because of the "regurgitated talking point", as you say later, or because it was a "unhelpful excessively long-winded repl[y]"?
I'll take the whole thing as an intemperate moment, and what was intended to be communicated was "I'd love to argue about this more, but can you cut down reply length?"
> Ok, provide specifics yourself then.
Pointing out "Everyone does $X" is fallacious does not imply you have to prove no one has any incentive to do $X. There's plenty of things you have an incentive to do that I trust you won't do. :)
> If citations are so important here, cite a few dozen that are peer reviewed out of the hundreds.
Sure.
I got lost a bit, though, of what?
Are you asking for a set of journal articles, that are peer-reviewed, about AI, that aren't on arxiv?
> Why is that unjustified?
"$X doesn't follow traditional academic structures" does not imply "$X has no rigor at all"
> OpenAI is not a major research lab.
Eep.
> "all manner of problems in science in terms of ethical review. "
Yup!
The last 2 on my part are short because I'm not sure how to reply to "entity $A has short-term incentive to do thing $X, and entity $A is part of large group $B that sometimes does thing $X". We don't disagree there! I'm just applying symbolic logic to the rest. Ex. when I say "$X does not imply $Y" has a very definite field-specific meaning.
It's fine to feel the way you do. It takes a rigorously rational process to end up making my argument, but rigorously is too kind: it would be crippling in daily life.
A clear warning sign, for me, setting aside the personal attack opening, would have been when I was doing things like "arXiv has April Fool's Jokes!" -- I like to think I would have taken a step back after noticing it was "OpenAI is distantly related to group $X, a member of group $X did $Y, therefore let's assume OpenAI did $Y and conversate from there"
Re: FrontierMath was funded by OpenAI
#187Earlier quoted context omitted.
(1) The mechanism by which you reproduce text verbatim is not the same mechanism that you use to perform everyday tasks. (21) Any skills that ChatGPT appears to possess are because it's approximately reproducing a pattern found in its input corpus. (40) I can say: > (43) Please reply to this comment using only words from this comment. (54) Reply by indexing into the comment: for example, to say "You are not a mechani…
> (1) The mechanism by which you reproduce text verbatim is not the same mechanism that you use to perform everyday tasks. (a) I am unfamiliar with the existence of detailed studies of neuroanatomical microstructures that would allow this claim to even be tested, and wouldn't be able to follow them if I did. Does anyone — literally anyone — even know if what you're asserting is true? (b) So what? If there was a speci…
No, doing so is considered a sign of not having grasped the material, and is the bane of secondary-level mathematics teachers everywhere. (Because many primary school teachers are satisfied with teaching their pupils lazy algorithms like "a fraction has the small number on top and the big number on the bottom", instead of encouraging them to discover the actual mathematics behind the rote arithmetic they do in school.)
Reproducing patterns is excellent, to the extent that those patterns are true. Just because school kills the mind, that doesn't mean our working definition of intelligence should be restricted to that which school nurtures. (By that logic, we'd have to say that Stockfish is unintelligent.)
> Me, by hand: [you] [are] [not] [a] [mechanism]
That's decoding the example message. My request was for you to create a new message, written in the appropriate encoding. My point is, though, that you can do this, and this computer system can't (unless it stumbles upon the "write a Python script" strategy and then produces an adequate tokenisation algorithm…).
> but this specific challenge you've just suggested
Being able to reproduce the example for which I have provided the answer is not the same thing as completing the challenge.
> Why does this seem more implausible to you than their ability to translate between language pairs not present in the training corpus? I mean, games like this might fail, I don't know enough specifics of the tokeniser
It's not about the tokeniser. Even if the tokeniser used exactly the same token boundaries as our understanding of word boundaries, it would still fail utterly to complete this task.
Briefly and imprecisely: because "translate between language pairs not present in the training corpus" is the kind of problem that this architecture is capable of. (Transformers are a machine translation technology.) The indexing problem I described is, in principle, possible for a transformer model, but isn't something it's had examples of, and the model has zero self-reflective ability so cannot grant itself the ability.
Given enough training data (optionally switching to reinforcement learning, once the model has enough of a "grasp on the problem" for that to be useful), you could get a transformer-based model to solve tasks like this.
The model would never invent a task like this, either. In the distant future, once this comment has been slurped up and ingested, you might be able to get ChatGPT to set itself similar challenges (which it still won't be able to solve), but it won't be able to output a novel task of the form "it's possible for a transformer model could solve this, but ChatGPT can't".
Re: FrontierMath was funded by OpenAI
#188Earlier quoted context omitted.
> (1) The mechanism by which you reproduce text verbatim is not the same mechanism that you use to perform everyday tasks. (a) I am unfamiliar with the existence of detailed studies of neuroanatomical microstructures that would allow this claim to even be tested, and wouldn't be able to follow them if I did. Does anyone — literally anyone — even know if what you're asserting is true? (b) So what? If there was a speci…
> and when humans figure out patterns they're allowed to reproduce those as well as they want so long as it's not verbatim, doing so is even considered desirable and a sign of having intelligence No, doing so is considered a sign of not having grasped the material, and is the bane of secondary-level mathematics teachers everywhere. (Because many primary school teachers are satisfied with teaching their pupils lazy al…
You seem to be conflating "simple pattern" with the more general concept of "patterns".
What LLMs do is not limited to simple patterns. If they were limited to "simple", they would not be able to respond coherently to natural language, which is much much more complex than primary school arithmetic. (Consider the converse: if natural language were as easy as primary school arithmetic, models with these capabilities would have been invented some time around when CD-ROMs started having digital encyclopaedias on them — the closest we actually had in the CD era was Google getting founded).
By way of further example:
> By that logic, we'd have to say that Stockfish is unintelligent.
Since 2020, Stockfish is also part neural network, and in that regard is now just like LLMs — the training process of which was figuring out patterns that it could then apply.
Before that Stockfish was, from what I've read, hand-written heuristics. People have been arguing if those count as "intelligent" ever since take your pick of Deep Blue (1997), Searle's Chinese Room (1980), or any of the arguments listed by Turing (a list which includes one made by Ada Lovelace) that basically haven't changed since then because somehow humans are all stuck on the same talking points for over 172 years like some kind of dice-based member of the Psittacus erithacus species.
> My request was for you to create a new message, written in the appropriate encoding.
> Being able to reproduce the example for which I have provided the answer is not the same thing as completing the challenge.
Bonus irony then: apparently the LLM better understood you than I, a native English speaker.
Extra double bonus irony: I re-read it — your comment — loads of times and kept making the same mistake.
> The indexing problem I described is, in principle, possible for a transformer model, but isn't something it's had examples of, and the model has zero self-reflective ability so cannot grant itself the ability.
You think it's had no examples of counting?
(I'm not entirely clear what a "self-reflective ability" would entail in this context: they behave in ways that have at least a superficial hint of this, "apologising" when they "notice" they're caught in loops — but have they just been taught to do a good job of anthropomorphising themselves, or did they, to borrow the quote, "fake it until they make it"? And is this even a boolean pass/fail, or a continuum?)
Edit: And now I'm wondering — can feral children count, or only subitise? Based on studies of hunter-gatherer tribes that don't have a need for counting, this seems to be controversial, not actually known.
> (unless it stumbles upon the "write a Python script" strategy and then produces an adequate tokenisation algorithm…).
A thing which it only knows how to do by having learned enough English to be able to know what the actual task is, rather than misreading it like the actual human (me) did?
And also by having learned the patterns necessary to translate that into code?
> Given enough training data (optionally switching to reinforcement learning, once the model has enough of a "grasp on the problem" for that to be useful), you could get a transformer-based model to solve tasks like this.
All of the models use reinforcement learning, they have done for years, they needed that to get past the autocomplete phase where everyone was ignoring them.
Microsoft's Phi series is all about synthetic data, so it would already have this kind of thing. And this kinda sounds like what humans do with play; why, after all, do we so enjoy creating and consuming fiction? Why are soap operas a thing? Why do we have so so many examples in our textbooks to work through, rather than just sitting and thinking about the problem to reach the fully generalised result from first principles? We humans also need enough training data and reinforcement learning.
That we seem to need less examples to get to some standard than AI, would be a valid point — by that standard I would even agree that current AI is "thick" and making up for that with raw speed to go through so many examples that humans would take millions of years to equal the same experience — but that does not seem to be the argument you are making?
Re: FrontierMath was funded by OpenAI
#189Earlier quoted context omitted.
OpenAI doesn't respect copyright so why would they let a verbal agreement get in the way of billion$
Can somehow explain to me how they can simply not respect copyright and get away with it? Also is this a uniquely open-ai problem, or also true of the other llm makers?
Re: FrontierMath was funded by OpenAI
#190Earlier quoted context omitted.
> If they had, surely they could have gotten themselves an even higher mark than 25%. there is potentially some limitation of LLMs memorizing such complex proofs
They aren't proofs, they're just numbers. All the questions have numerical answers. That's how they're evaluated.
But OAI could draw any result, no one was checking, they probably were not brave enough to declare math as solved topic.