Earlier quoted context omitted.
> the copyright holder’s tort would be with the source of the infringing distribution, not the people who read the material. Someone who just reads the material doesn't infringe. But someone who copies it, or prepares works that are derivative of it (which can happen even if they don't copy a single word or phrase literally), does. > would I then owe the authors of the books I learned from a fee to apply that knowled…
Right, but the onus of responsibility being on the end user publishing the song or creative work in violation of copyright, not the text editor, word processor, musical notation software, etc, correct? A text prediction tool isn’t a person, the data it is trained on is irrelevant to the copyright infringement perpetrated by the end user. They should perform due diligence to prevent liability.
FrontierMath was funded by OpenAI
151–160 of 212 posts
Re: FrontierMath was funded by OpenAI
#152“… we have a verbal agreement that these materials will not be used in model training” Ha ha ha. Even written agreements are routinely violated as long as the potential upside > downside, and all you have is verbal agreement? And you didn’t disclose this? At the time o3 was released I wrote “this is so impressive that it brings out the pessimist in me”[0], thinking perhaps they were routing API calls to human workers…
>perhaps they were routing API calls to human workers Honest question, did they?
Re: FrontierMath was funded by OpenAI
#153Earlier quoted context omitted.
> if they used it in training it should be 100% hit. Not necessarily, no. A statistical model will attempt to minimise overall loss, generally speaking. If it gets 100% accuracy on the training data it's usually an overfit. (Hugging the data points too tightly, thereby failing to predict real life cases)
you are mostly right. but seeing almost perfectly reconstructed images from training set it's obvious model -can- memorize samples. in this case it would reproduce the answers too close to the original to be just 'accidental'. should be easy to test. My guess samples could be used to find good enough stopping point for o1, o3 models. which is hardcoded.
Re: FrontierMath was funded by OpenAI
#154Earlier quoted context omitted.
>perhaps they were routing API calls to human workers Honest question, did they?
How would that even work? Aren’t the responses to the API equally fast as the Web interface? Can any human write a response with the speed of an LLM?
Re: FrontierMath was funded by OpenAI
#155Do people actually think OpenAI is gaming benchmarks? I know they have lost trust and credibility, especially on HN. But this is a company with a giant revenue opportunity to sell products that work. What works for enterprise is very different from “does it beat this benchmark”. No matter how nefarious you think sama is, everything points to “build intelligence as rapidly as possible” rather than “spin our wheels mes…
Re: FrontierMath was funded by OpenAI
#156Earlier quoted context omitted.
OpenAI doesn't respect copyright so why would they let a verbal agreement get in the way of billion$
Can somehow explain to me how they can simply not respect copyright and get away with it? Also is this a uniquely open-ai problem, or also true of the other llm makers?
Uber showed the way. They initially operated illegally in many cities but moved so quickly as to capture the market and then they would tell the city that they need to be worked with because people love their service.
https://www.theguardian.com/news/2022/jul/10/uber-files-leak...
Re: FrontierMath was funded by OpenAI
#157“… we have a verbal agreement that these materials will not be used in model training” Ha ha ha. Even written agreements are routinely violated as long as the potential upside > downside, and all you have is verbal agreement? And you didn’t disclose this? At the time o3 was released I wrote “this is so impressive that it brings out the pessimist in me”[0], thinking perhaps they were routing API calls to human workers…
OpenAI doesn't respect copyright so why would they let a verbal agreement get in the way of billion$
I’d prefer we go the other direction where something like archive.org archives all publicly accessible content and the government manages this, keeps it up-to-date, and gives cheap access to all of the data to anyone on request. That’s much more “democratizing” than further locking down training data to big companies.
Re: FrontierMath was funded by OpenAI
#158Earlier quoted context omitted.
Right, but the onus of responsibility being on the end user publishing the song or creative work in violation of copyright, not the text editor, word processor, musical notation software, etc, correct? A text prediction tool isn’t a person, the data it is trained on is irrelevant to the copyright infringement perpetrated by the end user. They should perform due diligence to prevent liability.
How is the end user the one doing the infringement though? If I chat with ChatGPT and tell it „give me the first chapter of book XYZ“ and it gives me the text of the first chapter, OpenAI is distributing a copyrighted work without permission.
Re: FrontierMath was funded by OpenAI
#159Earlier quoted context omitted.
How is the end user the one doing the infringement though? If I chat with ChatGPT and tell it „give me the first chapter of book XYZ“ and it gives me the text of the first chapter, OpenAI is distributing a copyrighted work without permission.
Can you do that though? Just ask ChatGPT to give you the first chapter of a book and it gives it to you?
Not a book chapter specifically but this could already be considered copyright infringement, I think.
Re: FrontierMath was funded by OpenAI
#160Earlier quoted context omitted.
Right, but the onus of responsibility being on the end user publishing the song or creative work in violation of copyright, not the text editor, word processor, musical notation software, etc, correct? A text prediction tool isn’t a person, the data it is trained on is irrelevant to the copyright infringement perpetrated by the end user. They should perform due diligence to prevent liability.
> A text prediction tool isn’t a person, the data it is trained on is irrelevant to the copyright infringement perpetrated by the end user. They should perform due diligence to prevent liability. Huh what? If a program "predicts" some data that is a derivative work of some copyrighted work (that the end user did not input), then ipso facto the tool itself is a derivative work of that copyrighted work, and illegal to…
You learned English, math, social studies, science, business, engineering, humanities, from a McGraw Hill textbook? Sorry, all creative works you’ve produced are derivative of your educational materials copyrighted by the authors and publisher.