Earlier quoted context omitted.
I disagree that LLM models are the product of enormous quantities of copyright infringement. The recent announcement that AI-assisted research produced a counterexample to the Jacobian conjecture--a long-standing open problem in algebraic geometry--shows the original value AI can create. The result was not copied from a textbook; it emerged from AI learning from existing material, much as a human does, and then apply…
If you re-read your comment, you will find that your second paragraph is not evidence for the claim you make in your first paragraph. In fact, your first paragraph is just false.
“We have information that Moonshot distilled Fable for the development of K3”
441–450 of 742 posts
Re: “We have information that Moonshot distilled Fable for the development of K3”
#442Does this matter? Distillation is not illegal by every definition of the word. There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them. Another example is that it appears that the upper limit of what you can do is ultimately dependent on people working on the model, otherwise grok would be a LOT more competitive pre…
Boy I tell you, I am having an awful hard time summoning pity for the organizations that have themselves distilled all of humanity's knowledge into mysterious labor-market-masticating black boxes.
> protect our first-party products from abuse like bots, scraping
Won't you think of the trillion dollar corporations?!
Re: “We have information that Moonshot distilled Fable for the development of K3”
#443Earlier quoted context omitted.
No - distillation is not data inputs. Raw materials vs. Value add. They are different things, like ore and metal. Distillation is a new thing we need to understand, it's probably closer to IP than not.
distillation has been around for 12 years. it's not new in terms of ML techniques. https://arxiv.org/abs/1503.02531 although i doubt there has been a legal case over it yet in the context of the legality of stealing shit but IANAL.
It's completey insane that we still don't know how Open Source would work, that the laws are vague and we're still technically waiting for the courts to decide on cases.
The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.
Re: “We have information that Moonshot distilled Fable for the development of K3”
#444Earlier quoted context omitted.
Plenty of HN readers feel this way and it's a good point, but it has also become an entirely cliché response which pops up like mushrooms anytime "distillation" appears. That means it's against the site guidelines, which ask: " Eschew flamebait. Avoid generic tangents. Omit internet tropes. " - https://news.ycombinator.com/newsguidelines.html I don't mean to pick on you personally! It's just that reflexive responses…
You just told on yourself big time about being a lapdog for the US Feds. Easily one of the worst moderation decisions you’ve ever made and that’s impressive given your track record.
Re: “We have information that Moonshot distilled Fable for the development of K3”
#445Re: “We have information that Moonshot distilled Fable for the development of K3”
#446Earlier quoted context omitted.
> /giphy nobody cares SpongeBob meme Can you please not do this here? There's nothing wrong with it, we're just trying for something else on this site. " Don't be snarky. [...] Omit internet tropes. [...etc...] https://news.ycombinator.com/newsguidelines.html
Genuinely want to know if you're surprised that there have been almost no commenters concurring with (what seems to be) a well-substantiated (& on the face of it defensibly center-right) tweet from the WH Will contrarian dynamic take off here? Will it be flagged? Will you unflag it? Edit: catigula, mattrighetti With more substantive technical comments towards the middle
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor...
Re: “We have information that Moonshot distilled Fable for the development of K3”
#447Earlier quoted context omitted.
My parents put in countless hours and tens of thousands of dollars into raising me to the point where I could write an answer on StackOverflow And OpenAI scraped and distilled that answer and gave me nothing
And now people such as myself have access to open weight models with that information. I wasn't lucky enough to have parents put me through school, and LLMs have absolutely helped me further educate myself and play "catch up" on opportunities others have been given. So, the net effect has been (and is continuing to be) a democratization of information.
Re: “We have information that Moonshot distilled Fable for the development of K3”
#448Earlier quoted context omitted.
> on the same level like Anthropic scraped copyright protected material for their training. I see no problem with distillation, on the other hand the complete dismissal of copyright by AI labs is pretty bad, I don’t think we should put them at the same level
> on the other hand the complete dismissal of copyright by AI labs Courts keep ruling over and over that an LLM trained on copyrighted works qualifies as a transformative work and is therefore fair use. They don't have to dismiss copyright law, this has always been allowed. The only thing they get in trouble for is pirating the works to get their hands on them.
*USA only.
the UK has fair dealing, which is more restrictive
https://www.gov.uk/guidance/exceptions-to-copyright#fair-dea...
https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...
Re: “We have information that Moonshot distilled Fable for the development of K3”
#449Earlier quoted context omitted.
This is a misrepresentation though. The LLM output, is not the same as the input - there is value add. Of course works used as raw inputs to LLMs required work and are reasonably subject to IP concerns - but they are different. It's possible that the LLM makers 'owe' the content creators that created the content they used to make their products - it's an interesting but separate question. We could very well end up wh…
> but they are different. How, and why? > We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible. That is the current state of legal rulings - LLM output is public domain, not copyrightable.
Our current laws simply weren’t built for this and I expect the legal status of LLM output is not going to be resolved until Congress actually legislates on this topic.
Re: “We have information that Moonshot distilled Fable for the development of K3”
#450Earlier quoted context omitted.
Moreover, reading a copyrighted book and learning from it is not theft.
Generating a set of weights is not learning.