Earlier quoted context omitted.
> it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity. you mean like the regulatory clarity surrounding stealing shit to make the LLMs in the first place? > [There is] extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we…
It's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists. It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ... ... but Chinese SOTA foundries directly using distillation as fair game. I don't th…
Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training.
You can't distill what you are not given - simple as that.
Are Chinese using the output of US models to help create some additional training data for their own in some way? Yes - quite possibly (e.g. LLM as judge), but its got nothing to do with distillation.