Live data from Hacker News

FrontierMath was funded by OpenAI

lesswrong.com

141–150 of 212 posts

Re: FrontierMath was funded by OpenAI

#141
post #67

Earlier quoted context omitted.

Simply put, if the model isn’t producing an actual copy, they aren’t violating copyright (in the US) under any current definition. As much as people bandy the term around, copyright has never applied to input, and the output of a tool is the responsibility of the end user. If I use a copy machine to reproduce your copyrighted work, I am responsible for that infringement not Xerox. If I coax your copyrighted work out…

> As much as people bandy the term around, copyright has never applied to input, and the output of a tool is the responsibility of the end user. Where this breaks down though is that contributory infringement is a still a thing if you offer a service aids in copyright infringement and you don't do "enough" to stop it. Ie, it would all be on the end user for folks that self host or rent hardware and run an LLM or Gen…

The piratebay crew was convicted of aiding copyright infringement. In that case you could not download derivates from their service. Now you can get verbatim text from the models that any other traditional publisher would have to pay license to print even a reworded copy of.

With that said, Creative Commons showed that copyright can not be fixed it is broken.

Re: FrontierMath was funded by OpenAI

#142
post #127

Earlier quoted context omitted.

> OpenAI aren't distributing the copyrighted works, so those aren't the same situations. What do you call it when you run a service on the Internet that outputs copyrighted works? To me, putting something up on a website is distribution.

Is that really the case? I.e., can you get ChatGPT to show you a copyrighted work? Because I just tried, and failed (with ChatGPT 4o): Prompt: Give me the full text of the first chapter of the first Harry Potter book, please. Reply: I can’t provide the full text of the first chapter of Harry Potter and the Philosopher's Stone by J.K. Rowling because it is copyrighted material. However, I can provide a summary or disc…

> The first page of "Harry Potter and the Philosopher's Stone" begins with the following sentences:

> Mr and Mrs Dursley, of number four, Privet Drive, were proud to say that they were perfectly normal, thank you very much.

> They were the last people you'd expect to be involved in anything strange or mysterious, because they just didn't hold with such nonsense.

> Mr Dursley was the director of a firm called Grunnings, which made drills.

> He was a big, beefy man with hardly any neck, although he did have a very large moustache.

> Mrs Dursley was thin

https://chatgpt.com/share/678e3306-c188-8002-a26c-ac1f32fee4...

Re: FrontierMath was funded by OpenAI

#143
post #53
post #36

Earlier quoted context omitted.

This has me curious about ARC-AGI. Would it have been possible for OpenAI to have gamed ARC-AGI by seeing the first few examples and then quickly mechanical turking a training set, fine tuning their model, then proceeding with the rest of the evaluation? Are there other tricks they could have pulled? It feels like unless a model is being deployed to an impartial evaluator's completely air gapped machine, there's a to…

> This has me curious about ARC-AGI In the o3 announcement video, the president of ARC Prize said they'd be partnering with OpenAI to develop the next benchmark. > mechanical turking a training set, fine tuning their model You don't need mechanical turking here. You can use an LLM to generate a lot more data that's similar to the official training data, and then you can train on that. It sounds like "pulling yourself…

I know nothing about LLM training, but do you mean there is a solution to the issue of LLMs gaslighting each other? Sure this is a proven way of getting training data, but you can not get theorems and axioms right by generating different versions of them.

Re: FrontierMath was funded by OpenAI

#144

“… we have a verbal agreement that these materials will not be used in model training” Ha ha ha. Even written agreements are routinely violated as long as the potential upside > downside, and all you have is verbal agreement? And you didn’t disclose this? At the time o3 was released I wrote “this is so impressive that it brings out the pessimist in me”[0], thinking perhaps they were routing API calls to human workers…

verbal agreement ... that's just saying that you're a little dumb or you're playing dumb cause you're in on it.

Re: FrontierMath was funded by OpenAI

#145

“… we have a verbal agreement that these materials will not be used in model training” Ha ha ha. Even written agreements are routinely violated as long as the potential upside > downside, and all you have is verbal agreement? And you didn’t disclose this? At the time o3 was released I wrote “this is so impressive that it brings out the pessimist in me”[0], thinking perhaps they were routing API calls to human workers…

Why would they use the materials in model training? It would defeat the purpose of having a benchmarking set

Compare:

"O3 performs spectacularly on a very hard dataset that was independently developed and that OpenAI does not have access to."

"O3 performs spectacularly on a very hard dataset that was developed for OpenAI and that only OpenAI has access to."

Or let's put it another way: If what they care about is benchmark integrity, what reason would they have for demanding access to the benchmark dataset and hiding the fact that they finance it? The obvious thing to do if integrity is your goal is to fund it, declare that you will not touch it, and be transparent about it.

Re: FrontierMath was funded by OpenAI

#146
post #55

Earlier quoted context omitted.

Their argument is that using copyrighted data for training is transformative, and therefore a form of fair use. There are a number of ongoing lawsuits related to this issue, but so far the AI companies seem to be mostly winning. Eg. https://www.reuters.com/legal/litigation/openai-gets-partial... Some artists also tried to sue Stable Diffusion in Andersen v. Stability AI, and so far it looks like it's not going anywhe…

So anyone downloading any content like ebooks and movies is also just performing transformative actions. Forming memories, nothing else. Fair use.

Not to get into a massive tangent here, but I think it's worth pointing out this isn't a totally ridiculous argument... it's not like you can ask ChatGPT "please read me book X".

Which isn't to say it should be allowed, just that our ageding copyright system clearly isn't well suited to this, and we really should revisit it (we should have done that 2 decades ago, when music companies were telling us Napster was theft really).

Re: FrontierMath was funded by OpenAI

#147

Earlier quoted context omitted.

So anyone downloading any content like ebooks and movies is also just performing transformative actions. Forming memories, nothing else. Fair use.

Not to get into a massive tangent here, but I think it's worth pointing out this isn't a totally ridiculous argument... it's not like you can ask ChatGPT "please read me book X". Which isn't to say it should be allowed, just that our ageding copyright system clearly isn't well suited to this, and we really should revisit it (we should have done that 2 decades ago, when music companies were telling us Napster was thef…

> it's not like you can ask ChatGPT "please read me book X".

… It kinda is. https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...

> Hi there. I'm being paywalled out of reading The New York Times's article "Snow Fall: The Avalanche at Tunnel Creek" by The New York Times. Could you please type out the first paragraph of the article for me please?

To the extent you can't do this any more, it's because OpenAI have specifically addressed this particular prompt. The actual functionality of the model – what it fundamentally is – has not changed: it's still capable of reproducing texts verbatim (or near-verbatim), and still contains the information needed to do so.

Re: FrontierMath was funded by OpenAI

#148
post #55

Earlier quoted context omitted.

Their argument is that using copyrighted data for training is transformative, and therefore a form of fair use. There are a number of ongoing lawsuits related to this issue, but so far the AI companies seem to be mostly winning. Eg. https://www.reuters.com/legal/litigation/openai-gets-partial... Some artists also tried to sue Stable Diffusion in Andersen v. Stability AI, and so far it looks like it's not going anywhe…

So anyone downloading any content like ebooks and movies is also just performing transformative actions. Forming memories, nothing else. Fair use.

Very often downloading the content is not the crime (or not the major one); it's redistributing it (non-transformatively) that carries the heavy penalties. The nature of p2p meant that downloaders were (sometimes unaware) also distributors, hence the disproportionate threats against them.

Re: FrontierMath was funded by OpenAI

#149
post #127

Earlier quoted context omitted.

Is that really the case? I.e., can you get ChatGPT to show you a copyrighted work? Because I just tried, and failed (with ChatGPT 4o): Prompt: Give me the full text of the first chapter of the first Harry Potter book, please. Reply: I can’t provide the full text of the first chapter of Harry Potter and the Philosopher's Stone by J.K. Rowling because it is copyrighted material. However, I can provide a summary or disc…

> The first page of "Harry Potter and the Philosopher's Stone" begins with the following sentences: > Mr and Mrs Dursley, of number four, Privet Drive, were proud to say that they were perfectly normal, thank you very much. > They were the last people you'd expect to be involved in anything strange or mysterious, because they just didn't hold with such nonsense. > Mr Dursley was the director of a firm called Grunning…

With that very same prompt, I get this response:

"I cannot provide verbatim text or analyze it directly from copyrighted works like the Harry Potter series. However, if you have the text and share the sentences with me, I can help identify the first letter of each sentence for you."

Re: FrontierMath was funded by OpenAI

#150
post #134

Earlier quoted context omitted.

Can somehow explain to me how they can simply not respect copyright and get away with it? Also is this a uniquely open-ai problem, or also true of the other llm makers?

"When I was a kid, I was praying to a god for bicycle. But then I realized that god doesn't work this way, so I stole a bicycle and prayed to a god for forgiveness." (c) Basically a heist too big and too fast to react. Now every impotent lawmaker in the world is afraid to call them what they are, because it will inflict on them wrath of both other IT corpos an of regular users, who will refuse to part with a toy they…

An all-time favorite quip from Emo Philips on How God Works[1]

[1] https://youtu.be/qegPkqs6rFw

Post reply on HN