Earlier quoted context omitted.
Making an LLM from raw data is value-add. Distillation is just value extract. It's soft, and I'm not sure what the answer should be ... but I think that there is a difference. I think we start by recognizing that ... and then try to figure it out from there. 'The Internet' may be a public good, maybe we make them pay a tax for that, but that's different than distillation.
> Making an LLM from raw data is value-add. > Distillation is just value extract. There is a value-add in selecting the valuable parts out of the garbage. And let's face it. Largest models contain a lot of garbage.
We ought to identify that and integrate that into our thinking.