Author fears pending/future copyright litigation against AI labs, proposes "Corpus Royalty"
The Private Capture of Public Genius
91–100 of 111 posts
Re: The Private Capture of Public Genius
#92> Frontier science looks different today. It's rooted in model weights and GPUs. It is flooded with token spend and agentic loops. It blooms in data centers. This seems like handwaving to me. Even the actual frontier science that is using ML (e.g. AlphaFold) isn't based on "token spend" or "agentic loops". I personally cannot think of a single example of frontier science that is rooted in LLMs. I am sure there are a…
Re: The Private Capture of Public Genius
#93This was interesting right up until "The fund pays every eligible American the same amount each year. " I'm in Australia. I've contributed my share of dirt to the delta. Why do I not get a share of this? I get that the frontier companies are (for the moment) US companies. But that's just corporate ownership, it's not what we're talking about. We're talking about compensating the people who wrote the training data for…
Why do you think a UN-hosted process would work? Has the UN done similarly intrusive things before and been effective at it? How do you account for Chinese frontier and open weights models, which are just months behind American models and subsidized by a sovereign state that does not share any of these premises about intellectual property? A lot of these posts seem subtextually premised on the idea that it's possible…
Even if the UN only has a very small amount of legitimacy and credibility… that’s still far far better than literally zero.
Re: The Private Capture of Public Genius
#94Earlier quoted context omitted.
Why do you think a UN-hosted process would work? Has the UN done similarly intrusive things before and been effective at it? How do you account for Chinese frontier and open weights models, which are just months behind American models and subsidized by a sovereign state that does not share any of these premises about intellectual property? A lot of these posts seem subtextually premised on the idea that it's possible…
This comment seems incoherent? Even if the UN only has a very small amount of legitimacy and credibility… that’s still far far better than literally zero.
Re: The Private Capture of Public Genius
#95Earlier quoted context omitted.
This comment seems incoherent? Even if the UN only has a very small amount of legitimacy and credibility… that’s still far far better than literally zero.
Well, I mean, the alternative is for a US-based process to be put into place, followed (or led) by an EU process. The premise of my comment is that no matter who hosts the process, China isn't going to comply, so it really just comes down to the US and Europe.
There are over 190 nations in the UN. Even if we assume a chunk have no credible claim whatsoever, that’s still well over a hundred.
Re: The Private Capture of Public Genius
#96Earlier quoted context omitted.
Well, I mean, the alternative is for a US-based process to be put into place, followed (or led) by an EU process. The premise of my comment is that no matter who hosts the process, China isn't going to comply, so it really just comes down to the US and Europe.
What? There are over 190 nations in the UN. Even if we assume a chunk have no credible claim whatsoever, that’s still well over a hundred.
Re: The Private Capture of Public Genius
#97Earlier quoted context omitted.
What? There are over 190 nations in the UN. Even if we assume a chunk have no credible claim whatsoever, that’s still well over a hundred.
There are not in fact 100 nations that will seriously contend for frontier models.
I’m talking about training data, and so is the parent.
Re: The Private Capture of Public Genius
#98Earlier quoted context omitted.
There are not in fact 100 nations that will seriously contend for frontier models.
You need to re-read the comment chain. Nobody claimed so many nations would contend for frontier models. I’m talking about training data, and so is the parent.
Re: The Private Capture of Public Genius
#99We cannot always want to capture only the (temporary) winners whenever we see a lucrative business and expect to share a free ride. I'd also assume that most of the revenue these AI labs are making is turned into depreciating fixed capital (hardware) and OPEX at this point. Why don't we capture Meta and Google as they allegedly take advantage of more publicly available information for profit? Let alone the truly valu…
I think Google's AI results are probably the prime example here. Those results are often quite good, and starve off the visits and hence revenue stream of the sites where the results are sourced from. Additionally, there's no way to robustly attribute that information anymore either; those links were already broken pre-AI by freeloading aggregators. So potential information producers can't afford to host their valuable information since they will pay proportionally to their information's value (as translated into bandwidth).
But it's tricky. If you charge AI companies proportional to the damage they do, then you need to assess that damage and you'll be caught in a no-win cat & mouse game where the AI companies outsource the damage and you try to track it back to them. They win if they bump their revenue, but they also win if they conceal the damage. (Just like companies can use shell companies, bankruptcy, and asset-only purchases to avoid Superfund responsibility.) If you charge proportional to revenue, then there aren't conflicting incentives; companies win by increasing revenue full stop. But I do agree that this shouldn't just apply to the big companies; distillation / model extraction (adversarial or not) should not be a way to avoid fees.
Though what really seemed off to me about the proposed solution was the use of the fees (to pay Americans, no less!) Perhaps upcoming essays will justify this more, but to me it seems like those fees should be applied directly to the health of the ecosystem. It should be used to fight pollution, in this case AI slop overtaking the Web. It should support the value creators, eg open source developers being overwhelmed by the AI tsunami. It should go towards serving and moderating online communities that create the very value that the LLMs are trained from. In theory, paying a whole bunch of Americans a trickle of blood money will end up going towards these purposes, but I'm very skeptical: first, the benefit will be very diluted. Second, it's more likely that the excess money will end up in Google or Amazon's pockets. The whole system is already set up to route any advantage to the big players. The payments are at least as likely to damage the ecosystem as they are to support it. They would just feed the capitalistic wolf and further reinforce the setup where monetizers win and value creators lose. If you could magically route the money to value creators, that'd be great. But you can't.
It'd be ok to pour money into a bucket if it had a small hole or two, but pouring it faster into the sieve we have now is not going to help anything.
Re: The Private Capture of Public Genius
#100It's been said a thousand times but the ability to produce copyrighted material is not copyright infringement. I'll say it again: You can produce copyrighted material all day long and it's not copyright infringement. If I draw The Simpsons on a piece of paper - whether or not I used AI to create it - it's not copyright infringement. Copyright infringement is if I tried to sell that Simpsons work as my own, putting up…
Aside from fair use, yes it is.
> If I draw The Simpsons on a piece of paper - whether or not I used AI to create it - it's not copyright infringement.
Only because of the fair use doctrine, which is limited. If it damages the market for the original, then a court could definitely declare it to be infringement.
> Copyright infringement is if I tried to sell that Simpsons work as my own, putting up for consumption and reaping the monetary benefits.
No. Copyright infringement requires neither sale nor misrepresenting a work as one's one. It also covers derivative works, not just perfect duplicates. Copyright covers reproduction and derivate works and distribution/performance -- you can get in trouble for just one of those. Taken strictly, that would be a horrible world, but fortunately the fair use doctrine weakens those quite a bit. On the flip side, if AI reproduction destroys the original markets -- as it is quite obviously doing right now -- then it's going to have a reckoning at some point, given how its legal status is completely based upon getting a free ride by claiming fair use.