Live data from Hacker News

“We have information that Moonshot distilled Fable for the development of K3”

twitter.com

571–580 of 742 posts

Re: “We have information that Moonshot distilled Fable for the development of K3”

#571

Earlier quoted context omitted.

> it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity. you mean like the regulatory clarity surrounding stealing shit to make the LLMs in the first place? > [There is] extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we…

It's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists. It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ... ... but Chinese SOTA foundries directly using distillation as fair game. I don't th…

You're being too charitable to Anthropic, and assuming that the way they are abusing the word "distillation" has some real meaning here. It doesn't.

Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training.

You can't distill what you are not given - simple as that.

Are Chinese using the output of US models to help create some additional training data for their own in some way? Yes - quite possibly (e.g. LLM as judge), but its got nothing to do with distillation.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#572

Earlier quoted context omitted.

> it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity. you mean like the regulatory clarity surrounding stealing shit to make the LLMs in the first place? > [There is] extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we…

It's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists. It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ... ... but Chinese SOTA foundries directly using distillation as fair game. I don't th…

> it's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists.

for the record, i've always been rabidly pro-copyright since i worked at a performing royalty organization (prs for music) circa 15 years ago, way before i joined hn.

i don't use llms for that reason.

> It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...

when i see a spade, i call it a spade. just because the US has utterly stupid copyright provisions that are wide open for abuse, i.e. fair use, doesn't mean abusing those provisions at scale is morally acceptable.

> ... but Chinese SOTA foundries directly using distillation as fair game.

two wrongs don't make a right, but the irony is at least something.

> I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.

the corpos can get fucked as far as i'm concerned.

> What is more reasonable: ... There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.

*only in the US.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#573
post #468

Earlier quoted context omitted.

That they should be able to find distillation 'attacks' if they had enough observability.

That's not @throwa356262's argument. @throwa356262 argument is that it is infeasible to distill and release a new frontier model in two weeks.

Ah I interpreted it as 'of course they can't stop distillation if they couldn't stop a model from escaping its sandbox'

I can see how there's a big leap there, but I agree somewhat. If they are aware these are happening and can detect it as it is happening why are they not stopping them? What do you do there? It'll be cat and mouse for a while. Thinking of reasons they wouldn't try and stop it is just a lot of speculation in my brain.

It's probably a way harder problem than I think it is, but they are aware of them now, so I assume they are going to get more aggressive about it.

Let's say then that they can't detect them near real time or even a bit after, maybe they do have a big observabilty gap that no one has solved adequately.

The speed which they add features I've needed for governance is pretty close to the speed I 'manually' write those for my company. To me personally we are all just going fast and breaking everything and not having enough time to set up safe environments. I'm sure it's in the backlog.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#574

Kimi K3 was released July 16, Fable ban was lifted on July 1 but access was still limited. How did Moonshot "distil" a huge model in such short time and still had time to run the benchmarks and do the usual release thingies? I think Anthropic is desperate to stop foreign competition and the administration is happy to help because they too are heavily invested in these companies

Even if they distilled this crappy politician should have no issue. Anthropic pirated whole ebook collection and millions of github repo with gpl license. We should do more distillation and figure out how to create faster leaner and better models.

But training was ruled fair use, just the way they got the copies was illegal.

Like distillation?

Re: “We have information that Moonshot distilled Fable for the development of K3”

#576

Earlier quoted context omitted.

And now people such as myself have access to open weight models with that information. I wasn't lucky enough to have parents put me through school, and LLMs have absolutely helped me further educate myself and play "catch up" on opportunities others have been given. So, the net effect has been (and is continuing to be) a democratization of information.

Not to be an arse, but didn't you have access to Stack Overflow with all questions/answers prior to LLMs?

Yes, but time is finite

Re: “We have information that Moonshot distilled Fable for the development of K3”

#577

Earlier quoted context omitted.

And now people such as myself have access to open weight models with that information. I wasn't lucky enough to have parents put me through school, and LLMs have absolutely helped me further educate myself and play "catch up" on opportunities others have been given. So, the net effect has been (and is continuing to be) a democratization of information.

Democratization of information, but Sam Altman gets a $100B net worth and I'm still broke :) I would prefer some sort of democratiziation of the money made from the democratization of information as well

Agreed, i wish that it would have done more than change who gets rich off rent-seeking behavior surrounding the knowledge that others created, instead it just consolidated that from many gatekeepers to a few.

I can at least take some measure of pleasure in the fact that it has generally lessened the roadblocks in gathering information. I am still displeased that there are any gatekeepers of humanity's combined knowledge

Re: “We have information that Moonshot distilled Fable for the development of K3”

#578
post #533

Earlier quoted context omitted.

so what? this has nothing to do with them taking from those that took before them. china censoring information they do not like is nothing new and seems irrelevant to this conversation

Fair enough you don't care, some people might want to know about the Uyghur happy camps, mass organ harvesting and such. In a world where information sources are only going to dwindle, it is not in anyone's interest to empower actors that will use these to manipulate perceptions

[deleted]

Re: “We have information that Moonshot distilled Fable for the development of K3”

#579
https://www.americanheritage.com/copywrong-short-history-lit...

Now, of course it's in the "creators" interest to prevent their biz models to breakdown due to piracy but it will be interesting to see how it will turn out. Its similar to any other form of piracy in internet age: you can't pretend to have global distribution and absolute global control at the same time.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#580

Earlier quoted context omitted.

It's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists. It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ... ... but Chinese SOTA foundries directly using distillation as fair game. I don't th…

You're being too charitable to Anthropic, and assuming that the way they are abusing the word "distillation" has some real meaning here. It doesn't. Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training. You can't distill what you are not given - simple as that.…

I see your 'fine point' but I don't think it holds - 'distillation' is a perfectly reasonable term to describe the process of creating outputs from one model to that expose key training element, to use in another model.

I think where the definition may be be invalid, is in the creation of 'unrelated data sets for training' models, for unrelated issues.

Creating training sets that mach a models core training, is definitely distillation, it does not have to expose the reasoning traces.

Synthesizing data for some arbitrary thing ... I'm not sure that would be the same thing.

It's hard to draw the line.

But the Chinese models are absolutely distilling - and would not be competitive without this distillation.

At the same time, there's a lot of real innovation and regular building going on at the same time over there.

Post reply on HN