It's about the narrative that "Chinese models are at Fable level". The truth (if correct) is the China continues to copy, and the proprietary US Models continue to lead the state of the art.
There is no K4 without Fable 6, GPT-6. That, matters.
511–520 of 742 posts
It's about the narrative that "Chinese models are at Fable level". The truth (if correct) is the China continues to copy, and the proprietary US Models continue to lead the state of the art.
There is no K4 without Fable 6, GPT-6. That, matters.
Does this matter? Distillation is not illegal by every definition of the word. There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them. Another example is that it appears that the upper limit of what you can do is ultimately dependent on people working on the model, otherwise grok would be a LOT more competitive pre…
I agree distillation isn't illegal; I also think Moonshot/Kimi is very impressive. But the more interesting question is whether labs like Moonshot can be a real competitor to OpenAI/Anthropic. If you can only play catchup (however quickly you do that), then you're never going to be at the frontier - I think that's why distillation matters.
The answer depends on whether you think the AI researchers at Chinese labs are (or can be) as smart, motivated, and as good at math as those working at US labs - a not-insignificant proportion of whom are Chinese nationals.
Earlier quoted context omitted.
Maybe, but it's not like their AI is likely to repeat it back verbatim so it's unlikely to be a copyright violation. It seems like at most, they would be breaking Anthropic's terms of service? Or maybe they're going through an intermediary "transfer station" that's breaking terms of service: https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...
Yes, it's just a ToS violation at present. Those are legally binding though, despite the common adage. What that really translates to here though, anyone's guess. Anthropic's own copyright infringement could apparently be forgiven for 1.5B USD after all, so maybe there's a price that breaking the distillation clause for is acceptable too. Or some other arrangement.
Why is Anthropic's ToS any more binding than that of a rabidly-anti-ai literature blog with 50 readers?
Earlier quoted context omitted.
> Anthropic agreed to pay $1.5B in a settlement to avoid losing a pretty cut and dry case. It takes two parties to agree to a settlement. That the other party agreed to a settlement instead of taking it to court implies this was not the slam dunk you may think it was.
You're both reading tea leaves. Settling just says that they expected the internal costs or risks to be more than 1.5 billion cashflow. In the $65B in Series H funding at $965B post-money valuation they said their run-rate revenue crossed $47B annualised. With those numbers, there can be sound financial reasons for wanting to just get rid of the lawsuit. Also if it ends up that other competitors also need to pay $1.5…
Yes, of course, on Anthropic's side. Why would the other side agree to a settlement?
Earlier quoted context omitted.
Yes, I get that, but it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity. It's completey insane that we still don't know how Open Source would work, that the laws are vague and we're still technically waiting for the courts to decide on cases. The government should a) legislate and b) create test cases and run them through the courts so that we can have c…
> it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity. you mean like the regulatory clarity surrounding stealing shit to make the LLMs in the first place? > [There is] extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we…
It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...
... but Chinese SOTA foundries directly using distillation as fair game.
I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.
What is more reasonable:
- There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.
- SOTA makers are producing novel works, there is value add in that process, again roughly speaking.
- Distillation is a bit of a grey zone, producing random content as arbitrary input is one thing, but producing training sets is another. I think there's a coherent line in there somewhere, I'm not sure where it is.
Earlier quoted context omitted.
>a level of creativity in model creation that isn't present in distillation. the same argument - a level of creativity in the world knowledge creation that ins't present in the model training on that knowledge. Or in other words - model creation and training is just a distilling of the world knowledge.
I don't disagree. I'm not sure why that's a relevant reply though. If you think that the addition of a less creative process (model creation) to a more creative corpus ("art") is problematic, then it follows that you should think the addition of a less creative process (distillation) to a more creative corpus (a model) is also problematic.
Earlier quoted context omitted.
Of course it matters. Regardless of whether distillation is legal, there is a difference between training a model with and without distillation. For one thing, the distilled model wouldn't exist without the model it distilled. Also, companies that use distillation may be competitive but seem unlikely to surpass the companies that are training these models from scratch.
> but seem unlikely to surpass the companies that are training these models from scratch Then why is it a problem? Another serious question. Trying to get my head around what the root of the objection is here. There must be some fear, but if that fear is not a fear of being surpassed in the market, then what is the fear?
Earlier quoted context omitted.
This is BS to pressure politicians. Even an openai's guy (head of something made up) called bs on the idea you can train something like k3 by distillation. Anybody I know who works in LLM research says that distillation is either useless or merely useful in post training to show "correct" behavior. And even then you don't get a competing model, if RL on good prompts was that useful, all labs would've long skyrocketed…
Dean Ball, "head of strategic futures" at openai. https://xcancel.com/deanwball/status/2078133895766114412#m
Point four is especially telling.
Ball is deeply terrified of "AI communism", or in less red-scarey terms a world where AI is a public good and him and his fellow oligarchs don't get to centralize the accumulated knowledge of all of humanity and charge rent for it.
I think he's so deeply stuck in his ideological bubble he can't conceive that what he describes as a dystopia is the only way the future wouldn't be a dystopia for the vast majority of people.
Or to put it more clearly, the oligarch utopia he's trying to build is dystopia for the vast majority of humanity. The "utopia" he's trying to build is one of riches for him and serfdom for us.
Earlier quoted context omitted.
Because model output is probably far closer to software or a licensed work which possibly has greater protections than it is to copyright. There is far less possibility of fair use, it might be protected by patents, license or reverse engineering laws. In any case the laws are being written now, but I doubt these will have worse protection than software does, which has far better protections than copyright
> I doubt these will have worse protection than software does, which has far better protections than copyright Software is protected by copyright. Some software may also be protected by patents, but last time I checked, AI generated output of any kind was not patentable.
Also note that the OpenAI/Anthropic argument is that the model training is sufficiently transformative to satisfy the fair use of the original content for training.
By that same argument, when distilling the distillers aren't using the original content the OpenAI/Anthropic models were trained on - the distillers are interacting only with the "sufficiently transformed" content of the OpenAI/Anthropic models and are normally paying for that.
There is also that old phonebook rule that facts can't be copyrighted. So, if i asked the model about bunch of phone numbers, i can publish the resulting list, can train my model on it, etc. Such approach doesn't allow to reproduce copyrighted works of course - and as we know the AI output isn't copyrightable, so it looks like basically any output i get i can use whatever way i like.
Kimi K3 was released July 16, Fable ban was lifted on July 1 but access was still limited. How did Moonshot "distil" a huge model in such short time and still had time to run the benchmarks and do the usual release thingies? I think Anthropic is desperate to stop foreign competition and the administration is happy to help because they too are heavily invested in these companies