Live data from Hacker News

The False Promise of Imitating Proprietary LLMs

arxiv.org

81–90 of 90 posts

Re: The False Promise of Imitating Proprietary LLMs

#81

Earlier quoted context omitted.

Was GPT-4 trained on data that was acquired illegally? Or was it trained on data acquired legally that OpenAI didn't have the rights to redistribute? There is a difference. In the latter case, whether it counts as "stealing" would come down to whether or not GPT-4 counts as a derivative work, or some similar legal concept.

https://www.washingtonpost.com/technology/interactive/2023/a... Scribd has lots of pdfs of books that are copyrighted. The Washington Post article mentions there are several other places it downloaded and scraped pdfs of copyrighted textbooks, etc

That's interesting to know, but that doesn't by itself imply that it's illegal. For example, Google Books, which has massive amounts of scanned PDFs of copyrighted works, is considered fair use under US copyright law.

Re: The False Promise of Imitating Proprietary LLMs

#82

The jump between llama 13B and 30B is quite significant. And their instruction finetuning is not SOTA I don't think, though the point about general knowledge is a good one: instruction llama lies very confidently. But one great thing about open source LLMs is that you can specialize them in various tasks with affordable LORA training, enough to easily beat GPT4 in a specific niche.

Any recommended starting points for LORA training llama 30B on a specific niche? Books, tutorials, videos are all appreciated. Thanks for your time!

Currently SOTA for specialization of LLMs is QLoRA: https://github.com/artidoro/qlora

Re: The False Promise of Imitating Proprietary LLMs

#83
post #82

Earlier quoted context omitted.

Any recommended starting points for LORA training llama 30B on a specific niche? Books, tutorials, videos are all appreciated. Thanks for your time!

Currently SOTA for specialization of LLMs is QLoRA: https://github.com/artidoro/qlora

Thank you very much my friend. Be blessed in all that you do. <3

Re: The False Promise of Imitating Proprietary LLMs

#84

Earlier quoted context omitted.

https://www.washingtonpost.com/technology/interactive/2023/a... Scribd has lots of pdfs of books that are copyrighted. The Washington Post article mentions there are several other places it downloaded and scraped pdfs of copyrighted textbooks, etc

That's interesting to know, but that doesn't by itself imply that it's illegal. For example, Google Books, which has massive amounts of scanned PDFs of copyrighted works, is considered fair use under US copyright law.

As long as you don't try to scrape all the book's content…

It's only fair use for search purposes.

Re: The False Promise of Imitating Proprietary LLMs

#85
post #56

Earlier quoted context omitted.

So, LLMs aligned with the interests of our corporate overlords and that nebulous "national security" thing that somehow always translates to more surveillance and less due process?

This tech has about as much chance to continue unregulated as highly enriched uranium. There is no future-path that includes unregulated AI. I don't like horrific government abuse of residents,and I would not mind throwing most billionaire CEOs into a pool of alligators and dissolving their corporations. I don't like Altman, I think he's a smart person with NOBUS-level reckless hubris who is softballing the magnitude…

Fewer people agree with your premise, and that’s fortunate.

The “AI is dangerous” premise has no basis whatsoever. No one can prove it. No one can present a great thought experiment. Just doomsaying coupled with volume.

It’s starting to come off like a hidden agenda.

Re: The False Promise of Imitating Proprietary LLMs

#86

Earlier quoted context omitted.

That's interesting to know, but that doesn't by itself imply that it's illegal. For example, Google Books, which has massive amounts of scanned PDFs of copyrighted works, is considered fair use under US copyright law.

As long as you don't try to scrape all the book's content… It's only fair use for search purposes.

It's fair use if the work is "transformative". GPT-4 isn't publishing the content of the books, it's publishing a model derived from the entire corpus. I'm not a lawyer, but I think there's an argument that it is transformative.

Re: The False Promise of Imitating Proprietary LLMs

#87

Earlier quoted context omitted.

https://www.washingtonpost.com/technology/interactive/2023/a... Scribd has lots of pdfs of books that are copyrighted. The Washington Post article mentions there are several other places it downloaded and scraped pdfs of copyrighted textbooks, etc

That's interesting to know, but that doesn't by itself imply that it's illegal. For example, Google Books, which has massive amounts of scanned PDFs of copyrighted works, is considered fair use under US copyright law.

There's no good faith world where OPENAI trained only on legally available works

The only valid arguments is whether their model or it's output is itself protected legally.

Re: The False Promise of Imitating Proprietary LLMs

#88

Earlier quoted context omitted.

As long as you don't try to scrape all the book's content… It's only fair use for search purposes.

It's fair use if the work is "transformative". GPT-4 isn't publishing the content of the books, it's publishing a model derived from the entire corpus. I'm not a lawyer, but I think there's an argument that it is transformative.

Imho as transformative as encoding a DVD as DivX…

It's correct that OpenAI isn't publishing any of the "stolen" content directly. But they "stole" it to make their service possible in the first place. Not distributing it themself doesn't make much difference than.

Re: The False Promise of Imitating Proprietary LLMs

#89
post #56

Earlier quoted context omitted.

So, LLMs aligned with the interests of our corporate overlords and that nebulous "national security" thing that somehow always translates to more surveillance and less due process?

This tech has about as much chance to continue unregulated as highly enriched uranium. There is no future-path that includes unregulated AI. I don't like horrific government abuse of residents,and I would not mind throwing most billionaire CEOs into a pool of alligators and dissolving their corporations. I don't like Altman, I think he's a smart person with NOBUS-level reckless hubris who is softballing the magnitude…

I've heard many people advancing this thesis (usually by exactly the people who would benefit from such regulation), but no cogent arguments for it. Why do you think modern large-scale statistics needs to be regulated?

Re: The False Promise of Imitating Proprietary LLMs

#90

Earlier quoted context omitted.

This tech has about as much chance to continue unregulated as highly enriched uranium. There is no future-path that includes unregulated AI. I don't like horrific government abuse of residents,and I would not mind throwing most billionaire CEOs into a pool of alligators and dissolving their corporations. I don't like Altman, I think he's a smart person with NOBUS-level reckless hubris who is softballing the magnitude…

Fewer people agree with your premise, and that’s fortunate. The “AI is dangerous” premise has no basis whatsoever. No one can prove it. No one can present a great thought experiment. Just doomsaying coupled with volume. It’s starting to come off like a hidden agenda.

> Fewer people agree with your premise, and that’s fortunate.

Datacenter NVIDIA cards are already on the export control list for potential military use, and that was pre ChatGPT and GPT-4:

>On August 26, 2022, the U.S. government, or USG, informed NVIDIA Corporation, or the Company, that the USG has imposed a new license requirement, effective immediately, for any future export to China (including Hong Kong) and Russia of the Company’s A100 and forthcoming H100 integrated circuits. DGX or any other systems which incorporate A100 or H100 integrated circuits and the A100X are also covered by the new license requirement. The license requirement also includes any future NVIDIA integrated circuit achieving both peak performance and chip-to-chip I/O performance equal to or greater than thresholds that are roughly equivalent to the A100, as well as any system that includes those circuits. A license is required to export technology to support or develop covered products. The USG indicated that the new license requirement will address the risk that the covered products may be used in, or diverted to, a ‘military end use’ or ‘military end user’ in China and Russia. The Company does not sell products to customers in Russia.

https://www.sec.gov/Archives/edgar/data/1045810/000104581022...

> The “AI is dangerous” premise has no basis whatsoever. No one can prove it. No one can present a great thought experiment. Just doomsaying coupled with volume.

If you increase the number of persuasive Gobbels and hackers attacking infrastructure by 100,000x you do not come away with a better world.

> It’s starting to come off like a hidden agenda.

AI was used to fake the moon landing and hide bigfoot /s

Post reply on HN