On the legal front, I’ve been working with counsel to draft a counterclaim to Meta’s DMCA against llama-dl. (GPT-4 is surprisingly capable, but I’m talking to a few attorneys: https://twitter.com/theshawwn/status/1641841064800600070?s=6... ) An anonymous HN user named L pledged $200k for llama-dl’s legal defense: https://twitter.com/theshawwn/status/1641804013791215619?s=6... This may not seem like much vs Meta, but…
All models trained on public data need to be made public. As it is their outputs are not copyrightable, it’s not a stretch to say models are public domain.
If I create a website for tracking real estate trends in my area — which is public information — should I not be able to sell that information?
Similarly if a consulting company analyzes public market macro trends are they not allowed to sell that information?
Just because the information which is being aggregated and organized is public does not necessarily mean that the output product should be in the public.