Live data from Hacker News

“We have information that Moonshot distilled Fable for the development of K3”

twitter.com

451–460 of 742 posts

Re: “We have information that Moonshot distilled Fable for the development of K3”

#451
I think at some point the issues around copyrighted work and model distillation have to be disconnected to advance either idea.

1) Compensation of right holders is one issue.

2) Distilling models is an entirely separate issue, because model building is value add, and that is important because if we arrive at a place where you can produce a model, that gets to ~100% of what people perceive of the models value (on top of also not compensating right holders, yourself) you are discouraging development of better models and, again, in no way helping with issue 1)

Unless anyone actually distills a model and then also does something for rights holders, any schadenfreude simply detracts from this issue, in addition to the other issue (well, that might not be an issue if we would rather slow down model development right now, but again, forever worse models still don't help solve issue 1)

Re: “We have information that Moonshot distilled Fable for the development of K3”

#452

Commenters are overlooking the significance of this information and posting emotional reactions based on perceptions of fairness or feelings of schadenfreude. The economic viability of Anthropic and OpenAI rely on their being able to charge more for model access than their R&D and inference costs. If the market price for SOTA model access drops below that level, then these businesses will have to decide whether to co…

AI companies do not get to play the "Making an LLM using our data is unethical because the resulting LLM will replace us and hurt our profits" card.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#453

Earlier quoted context omitted.

If you re-read your comment, you will find that your second paragraph is not evidence for the claim you make in your first paragraph. In fact, your first paragraph is just false.

Let me explain it this way: If it is legal for a human to learn from a book, then disseminate the knowledge, then it is legal for a machine to do so. You may think this is not right because a machine does it at a much larger scale, but if so laws need to be updated. As it stands now there is no law that says if a human does X it is not a copyright violation but if a machine does the same X it is copyright violation.

> If it is legal for a human to learn from a book...

True, if the human's access to the book was legal

A great deal of training was on the open web, no one should complain.

But at least Meta and Anthropic were caught red handed taking copyrighted works, illegally, for training

I think international IP laws are too strick and onerous, but they were broken to train these models

Re: “We have information that Moonshot distilled Fable for the development of K3”

#454

Kimi K3 was released July 16, Fable ban was lifted on July 1 but access was still limited. How did Moonshot "distil" a huge model in such short time and still had time to run the benchmarks and do the usual release thingies? I think Anthropic is desperate to stop foreign competition and the administration is happy to help because they too are heavily invested in these companies

Do you get a token trophy for a few (many) trillion tokens purchased in distilation?

Re: “We have information that Moonshot distilled Fable for the development of K3”

#455

Earlier quoted context omitted.

I disagree that LLM models are the product of enormous quantities of copyright infringement. The recent announcement that AI-assisted research produced a counterexample to the Jacobian conjecture--a long-standing open problem in algebraic geometry--shows the original value AI can create. The result was not copied from a textbook; it emerged from AI learning from existing material, much as a human does, and then apply…

The didn’t pay for the books. It’s massive copyright infringement. The human buys the books.

It's worth keeping in mind the purpose of copyright. It's a pragmatic tool to encourage investment in creative work for the benefit of everybody/consumers. We may be entering a time where there's less need to incentivize people to write books. At least not non-fiction books which are simply a collection of existing knowledge presented in an a way that's suitable for human readers. A lot of the value those authors provided can now be done by AI. Yes, the AI trained on their work, but now that it's here, we don't need new non-fiction authors quite as much as we used to.

I wouldn't want to live in a world where technology or general people's wellbeing was held back by obsolete laws that ended up lingering on just to protect undeserving special people at the expense of the rest of society. Remember guilds for tradesmen? They were also a monopoly given by the government to special people. They had their purpose but nowadays we have different ways to keep tradesmen working effectively like license requirements and insurance.

Just to be clear, I think we do still need copyright, but that we might be in a transition period where it has to be redesigned to adapt to AI.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#458
post #278

Earlier quoted context omitted.

You can argue that reverse engineering anything is as hard if not harder than engineering something. I can’t imagine distillation is any different.

Distillation is objectively easier than training a model from scratch, that's why all these Chinese labs are doing it.

Training a model is objectively easier than generating the sum total of human creative output prior to 2020. That's why the big labs are doing it. What's the difference here?

Re: “We have information that Moonshot distilled Fable for the development of K3”

#459
post #137
post #90

Earlier quoted context omitted.

> on the same level like Anthropic scraped copyright protected material for their training. I see no problem with distillation, on the other hand the complete dismissal of copyright by AI labs is pretty bad, I don’t think we should put them at the same level

> on the other hand the complete dismissal of copyright by AI labs Courts keep ruling over and over that an LLM trained on copyrighted works qualifies as a transformative work and is therefore fair use. They don't have to dismiss copyright law, this has always been allowed. The only thing they get in trouble for is pirating the works to get their hands on them.

Man that's depressing to read someone defending this

Re: “We have information that Moonshot distilled Fable for the development of K3”

#460
post #336

Earlier quoted context omitted.

Not to mention the cases where the AI labs would have lost in court so bailed and settled for billions. Just this week, Anthropic agreed to pay $1.5B in a settlement to avoid losing a pretty cut and dry case.

> Anthropic agreed to pay $1.5B in a settlement to avoid losing a pretty cut and dry case. It takes two parties to agree to a settlement. That the other party agreed to a settlement instead of taking it to court implies this was not the slam dunk you may think it was.

You're both reading tea leaves.

Settling just says that they expected the internal costs or risks to be more than 1.5 billion cashflow.

  In the $65B in Series H funding at $965B post-money valuation they said their run-rate revenue crossed $47B annualised.
With those numbers, there can be sound financial reasons for wanting to just get rid of the lawsuit.

Also if it ends up that other competitors also need to pay $1.5 billion, then maybe that does or doesn't have a competitive advantage.

Anthropic's business and legal strategies are not public. I would expect there to be multiple legs/reasons for settlement even for a decision below 1%. Trying to create a single narrative is what us spectators do.

Post reply on HN