Live data from Hacker News

OpenAI now tries to hide that ChatGPT was trained on copyrighted books

businessinsider.com

1–10 of 94 posts

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#2
Unsurprising. Even worse than the Getty situation since OpenAI knew that they trained on copyrighted books for ChatGPT without the permission from the authors rather than getting a license just like they did with Shutterstock for DALLE-2.

The AI grift continues. Copying the same Silicon Valley playbook with a new narrative for promoting their AI snake-oil.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#3
I think the conclusion is slightly wrong: they're not trying to hide the training process. They'll probably wind up vigorously defending that in court however they can, they're in big trouble if they can't train like that.

They're trying to avoid reproducing copyrighted text, which is a totally separate (and arguably more clear-cut) legal question. Input vs output.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#4
post #2

Unsurprising. Even worse than the Getty situation since OpenAI knew that they trained on copyrighted books for ChatGPT without the permission from the authors rather than getting a license just like they did with Shutterstock for DALLE-2. The AI grift continues. Copying the same Silicon Valley playbook with a new narrative for promoting their AI snake-oil.

I'm not immediately recalling the Getty situation, could you perhaps enlighten me?

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#5
post #3

I think the conclusion is slightly wrong: they're not trying to hide the training process. They'll probably wind up vigorously defending that in court however they can, they're in big trouble if they can't train like that. They're trying to avoid reproducing copyrighted text, which is a totally separate (and arguably more clear-cut) legal question. Input vs output.

> Input vs output.

This is what I've been wondering. Does Fair Use apply here at all? Sure, the models were trained on copyrighted material. But wouldn't the generative part of the AI count as transformative?

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#6
I don't get it. Isn't that what people wanted to happen? Don't produce copyrighted work in the output. Sure you can learn from it, much like a director might learn from hundreds of movies he's watched. He obviously can't copy the plot from Die Hard but he can use elements he's picked up from it to make a Christmas movie.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#7
post #6

I don't get it. Isn't that what people wanted to happen? Don't produce copyrighted work in the output. Sure you can learn from it, much like a director might learn from hundreds of movies he's watched. He obviously can't copy the plot from Die Hard but he can use elements he's picked up from it to make a Christmas movie.

What do you mean he can't copy the plot? Any commercially successful movie is using one of about five basic plots.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#8
post #4
post #2

Unsurprising. Even worse than the Getty situation since OpenAI knew that they trained on copyrighted books for ChatGPT without the permission from the authors rather than getting a license just like they did with Shutterstock for DALLE-2. The AI grift continues. Copying the same Silicon Valley playbook with a new narrative for promoting their AI snake-oil.

I'm not immediately recalling the Getty situation, could you perhaps enlighten me?

AI-generated images were re-creating the Getty Images watermark (a transparent grey rectangle).
Post reply on HN