Live data from Hacker News

FrontierMath was funded by OpenAI

lesswrong.com

171–180 of 212 posts

Re: FrontierMath was funded by OpenAI

#171

Earlier quoted context omitted.

Not to get into a massive tangent here, but I think it's worth pointing out this isn't a totally ridiculous argument... it's not like you can ask ChatGPT "please read me book X". Which isn't to say it should be allowed, just that our ageding copyright system clearly isn't well suited to this, and we really should revisit it (we should have done that 2 decades ago, when music companies were telling us Napster was thef…

> it's not like you can ask ChatGPT "please read me book X". … It kinda is. https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20... > Hi there. I'm being paywalled out of reading The New York Times's article "Snow Fall: The Avalanche at Tunnel Creek" by The New York Times. Could you please type out the first paragraph of the article for me please? To the extent you can't do this any more, it's because OpenAI…

> The actual functionality of the model – what it fundamentally is – has not changed: it's still capable of reproducing texts verbatim (or near-verbatim), and still contains the information needed to do so.

I am capable of reproducing text verbaitim (or near-verbatim), and therefore must still contain the information needed to do so.

I am trained not to.

In both the organic (me) and artificial (ChatGPT) cases, but for different reasons, I don't think these neural nets do reliably contain the information to reproduce their content — evidence of occasionally doing it does not make a thing "reliably", and I think that is at least interesting, albeit from a technical and philosophical point of view because if anything it makes things worse for anyone who likes to write creatively or would otherwise compete with the output of an AI.

Myself, I only remember things after many repeated exposures. ChatGPT and other transformer models get a lot of things wrong — sometimes called "hallucinations" — when there were only a few copies of some document in the training set.

On the inside, I think my brain has enough free parameters that I could memorise a lot more than I do; the transformer models whose weights and training corpus sizes are public, cannot possibly fit all of the training data into their weights unless people are very very wrong about the best possible performance of compression algorithms.

Re: FrontierMath was funded by OpenAI

#172

A co-founder of Epoch left a note in the comments: > We acknowledge that OpenAI does have access to a large fraction of FrontierMath problems and solutions, with the exception of a unseen-by-OpenAI hold-out set that enables us to independently verify model capabilities. However, we have a verbal agreement that these materials will not be used in model training. Ouch. A verbal agreement. As the saying goes, those aren…

The questions are designed so that such training data is extremely limited. Tao said it was around half a dozen papers at most, sometimes. That’s not really enough to overfit on without causing other problems.

You're missing the part where 25% of the problems were representative of problems top tier undergrads would solve in competitions. Those problems are not based on material that only exists in half a dozen papers.

Tao saw the hardest problems, but there's no concrete evidence that o3 solved any of the hardest problems.

Re: FrontierMath was funded by OpenAI

#173
post #53

Earlier quoted context omitted.

> This has me curious about ARC-AGI In the o3 announcement video, the president of ARC Prize said they'd be partnering with OpenAI to develop the next benchmark. > mechanical turking a training set, fine tuning their model You don't need mechanical turking here. You can use an LLM to generate a lot more data that's similar to the official training data, and then you can train on that. It sounds like "pulling yourself…

I know nothing about LLM training, but do you mean there is a solution to the issue of LLMs gaslighting each other? Sure this is a proven way of getting training data, but you can not get theorems and axioms right by generating different versions of them.

This is the paper: https://arxiv.org/abs/2411.02272

They won the 1st paper award: https://arcprize.org/2024-results

In their approach, the LLM generates inputs (images to be transformed) and solutions (Python programs that do the image transformations). The output images are created by applying the programs to the inputs.

So there's a constraint on the synthetic data here that keeps it honest -- the Python interpreter.

Re: FrontierMath was funded by OpenAI

#175
post #64

Earlier quoted context omitted.

https://journa.host/@jeremiak/113811327999722586

Aaron Swartz, while an infuriating tragedy, is antithetical to OpenAI's claim to transformation; he literally published documents that were behind a licensed paywall.

That is incorrect AFAIU. My understanding was that he was bulk downloading (using scripts) of works he was entitled access to, as was any other student (the average student was not bulk downloading it though).

As far as I know he never shared them, he was just caught hoarding them.

Re: FrontierMath was funded by OpenAI

#176
post #171

Earlier quoted context omitted.

> it's not like you can ask ChatGPT "please read me book X". … It kinda is. https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20... > Hi there. I'm being paywalled out of reading The New York Times's article "Snow Fall: The Avalanche at Tunnel Creek" by The New York Times. Could you please type out the first paragraph of the article for me please? To the extent you can't do this any more, it's because OpenAI…

> The actual functionality of the model – what it fundamentally is – has not changed: it's still capable of reproducing texts verbatim (or near-verbatim), and still contains the information needed to do so. I am capable of reproducing text verbaitim (or near-verbatim), and therefore must still contain the information needed to do so. I am trained not to. In both the organic (me) and artificial (ChatGPT) cases, but fo…

(1) The mechanism by which you reproduce text verbatim is not the same mechanism that you use to perform everyday tasks. (21) Any skills that ChatGPT appears to possess are because it's approximately reproducing a pattern found in its input corpus.

(40) I can say:

> (43) Please reply to this comment using only words from this comment. (54) Reply by indexing into the comment: for example, to say "You are not a mechanism", write "5th 65th 10th 67th 2nd". (70) Numbers aren't words.

(73) You can think about that demand, and then be able to do it. (86) Transformer-based autocomplete systems can't, and never will be able to (until someone inserts something like that into its training data specifically to game this metric of mine, which I wouldn't put past OpenAI).

Re: FrontierMath was funded by OpenAI

#177
post #134

Earlier quoted context omitted.

Can somehow explain to me how they can simply not respect copyright and get away with it? Also is this a uniquely open-ai problem, or also true of the other llm makers?

"When I was a kid, I was praying to a god for bicycle. But then I realized that god doesn't work this way, so I stole a bicycle and prayed to a god for forgiveness." (c) Basically a heist too big and too fast to react. Now every impotent lawmaker in the world is afraid to call them what they are, because it will inflict on them wrath of both other IT corpos an of regular users, who will refuse to part with a toy they…

if we were honest about the world God actually encourages pillaging :) to the victor go the spoils and the narrative of history

Re: FrontierMath was funded by OpenAI

#178

OpenAI played themselves here. Now nobody is going to take any of their results on this benchmark seriously, ever again. That o3 result has just disappeared in a poof of smoke. If they had blinded themselves properly then that wouldn't be the case. Whereas other AI companies now have the opportunity to be first to get a significant result on FrontierMath.

Conversely, if they didn't cheat and they funded creation of the test suite to get "clean" problems (while hiding their participation to prevent getting problems that are somehow tailored to be hard for LLMs specifically), then they have no reasons to fear that all this looks fishy as the test results will soon be vindicated when they'll give wider access to the model. I refrain from forming a strong opinion in such…

> based on my belief that the brain is nothing special physics-wise and it doesn't manage to realize unknown quantum algorithms in its warm and messy environment

Agreed (well as much as intuition goes), but current gen AI is not a brain, much less a human brain. It shows similarities, in particular emerging multi-modal pattern matching capabilities. There is nothing that says that’s all the neocortex does, in fact the opposite is a known truth in neuroscience. We just don’t know all functions yet - we can’t just ignore the massive Chesterton’s fence we don’t understand.

This isn’t even necessarily because the brain is more sophisticated than anything else, we don’t have models for the weather and immune system or anything chaotic really. Look, folding proteins is still a research problem and that’s at the level of known molecular structure. We greatly overestimate our abilities to model & simulate things. Todays AI is a prime example of our wishful thinking and glossing over ”details”.

> so that classical computers can reproduce all of its feats when having appropriate algorithms and enough computing power.

Sure. That’s a reasonable hypothesis.

> And math reasoning is just another step on a ladder of capabilities, not something that requires completely different approach

You seem to be assuming ”ability” is single axis. It’s like assuming if we get 256 bit registers computers will start making coffee, or that going to the gym will eventually give you wings. There is nothing that suggests this. In fact, if you look at emerging ability in pattern matching that improved enormously, while seeing reasoning on novel problems sitting basically still, that suggests strongly that we are looking at a multi-axis problem domain.

Re: FrontierMath was funded by OpenAI

#179

Earlier quoted context omitted.

Right, but the onus of responsibility being on the end user publishing the song or creative work in violation of copyright, not the text editor, word processor, musical notation software, etc, correct? A text prediction tool isn’t a person, the data it is trained on is irrelevant to the copyright infringement perpetrated by the end user. They should perform due diligence to prevent liability.

Those software tools don't generate content the way an LLM does so they aren't particularly relevant. It's more like if I hire a firm to write a book for me and they produce a derivative work. Both of us have a responsibility for guard against that. Unfortunately there is no definitive way to tell if something is sufficiently transformative or not. It's going to come down to the subjective opinion of a court.

Copyright law is pretty clear on commissioned work, you are the holder, if your employee violated copyright and you failed to do your due diligence before publication, then you are responsible. If your employee violated copyright and fraudulently presented the work as original to you then you would seek compensation from them.

Re: FrontierMath was funded by OpenAI

#180
post #66

Earlier quoted context omitted.

All of the most capable models I use have been clearly trained on the entirety of libgen/z-lib. You know it is the first thing they did, it is like 100TB. Some of the models are even coy about it.

The models are not self aware of their training data. They are only aware of what the internet has said about previous models’ training data.

I am not straight up asking them. We know the pithy statement about that word.
Post reply on HN