Earlier quoted context omitted.
The entire case settled, the authors aren’t going to appeal when the company can’t hand out much more than the 1.5 billion in question and the company isn’t allowed to use the works in question going forward.
Before the settlement was made, Judge Alsup found as a matter of law that the training stage constituted fair use. https://storage.courtlistener.com/recap/gov.uscourts.cand.43...
California governor signs AI transparency bill into law
211–220 of 232 posts
Re: California governor signs AI transparency bill into law
#212Earlier quoted context omitted.
28yrs for an appeals chain is a bit longer than most realities I'm aware of. More like a dozen years at the top end would be more in line with what I've seen out there. In general though, it's easier to just comply, even for the companies. It helps with PR and employee retention, etc. They may fudge the reports a bit, even on purpose, but all groups of people do this to some degree. The question is, when does fudging…
For sure, it's meant to go a bit too far I guess. I'll be more real. This is to allow companies to make entirely fictitious statements that they will state satisfies their interpretation. The lack of fines will suggest compliance. Proving the statement is fiction isn't ever going to happen anyway. But it's also such low fine, that will eat inflation for those 12 years.
But I also understand that you were using hyperbole to emphasize your point, so there's not actually a reason to argue this.
Re: California governor signs AI transparency bill into law
#213Reading the text it feels like a giveaway to an "AI safety" industry who will be paid well to certify compliance.
The commenter’s profile indicates they work for a major AI development companies — where being against AI regulation aligns nicely with one’s paycheck. See also the the scare quotes around AI safety. We all have heard the dogma: regulation kills innovation. As if unbridled innovation is all that people and society care about. I wonder if the commenter above has ever worked in an industry where a safety culture matter…
I’m happy to stray outside the herd. HN needs more clearly articulated disagreement regarding AI regulation. I made my comment in response to what seemed like a simplistic, ideology-driven claim. Few wise hackers would make an analogous claim about a system they actually worked on. Thinking carefully about tech but phoning it in for other topics is a double standard.
Bare downvotes don't indicate (much less explain) one's rationale. I can’t tell if I (1) struck a nerve (an emotional response); (2) conjured a contentious philosophy (i.e. we have a difference in values or preferences or priorities), (3) made a logical error, (4) broke some norm or expectation, or (5) something else. Such a conflation of downvotes pushes me away from HN for meaningful discussion.
I’ve lived/worked on both coasts, Austin, and more, and worked at many places (startups, academic projects, research labs, gov't, not-for-profits, big tech) and I don’t consider myself defined by any one place or culture. But for the context of AI regulation, I have more fundamental priorities than anything close to "technical innovation at all costs".
P.S. (1) If a downvote here is merely an expression of “I’m a techno-libertarian” or "how dare you read someone's HN profile page and state the obvious?" or any such shallow disagreement, then IMO that’s counterproductive. If you want to express your viewpoint, make it persuasive rather than vaguely dismissive with an anonymous, unexplained downvote. (2) Some people do the thing where they guess at why someone else downvoted. That’s often speculation.
Re: California governor signs AI transparency bill into law
#214Earlier quoted context omitted.
Ostensibly copyright is there to increase economic incentives to make things it protects, and like you said, we can massively cut it down without affecting much there. So focusing on economic viability, set it to something like 15 years for code and 20-30 for everything else. Require registration for everything and source escrow for code and digital art to be granted copyright. That would give a wealth of code to tra…
Exclusively training on 15 year source code would make code generation significantly less useful as API’s change. Economic viability and utility for AI training are closely linked. Exclude all written works including news articles etc from the last 25 years and your model will know nothing about Facebook etc. It’s not as bad if you can exclude stuff from copyright and then use that, but your proposal would have obvio…
I suppose we all exist in our own bubbles, but I don't know why anyone would need a model that knows about Facebook etc. In any case, it's not clear that you couldn't train on news articles? AFAIK currently the only legal gray area with training is when e.g. Facebook mass pirated a bunch of textbooks. If you legally acquire the material, fitting a statistical model to it seems unlikely to run afoul of copyright law. Even without news articles, it would certainly learn something of the existence of Facebook. e.g. we are discussing it here, and as far as I know you're free to use the Hacker News BigQuery dump to your liking. Or in my proposed world, comments would naturally not be copyrighted since no one would bother to register them (and indeed a nominal fee could be charged to really make it pointless to do so). I suppose it is an important point that in addition to registration, we should again require notices, maybe including a registration ID.
Give a post-facto grace period of a couple weeks/months to register a thing for copyright. This would let you cover any work in progress that gets leaked by registering it immediately, causing the leak to become illegal.
Re: California governor signs AI transparency bill into law
#215Earlier quoted context omitted.
This is, like, the stupidest and most inefficient way for a government to make money. I know it's fun and all to circle jerk about how greedy those darn bureaucrats are - but we're all aware they control the budget, right? They could just raise taxes. I don't think they're fining companies... sigh... 10,000 dollars as some sort of sneaky "haha gotcha!" scam they're running.
Raising taxes does not solve the problem of getting the money out to the right people.
Re: California governor signs AI transparency bill into law
#216Earlier quoted context omitted.
Before the settlement was made, Judge Alsup found as a matter of law that the training stage constituted fair use. https://storage.courtlistener.com/recap/gov.uscourts.cand.43...
A judge yes, but that’s subject to appeal. The point is it never reached that stage and never will.
His opinion, while interlocutory and not binding precedent, will be cited in future cases. And his wasn't the only one. In Kadrey v. Meta Platforms, Inc., No. 23-cv-03417 (N.D. Cal. June 25, 2025) Judge Chhabria reached the same conclusion. https://storage.courtlistener.com/recap/gov.uscourts.cand.41...
In neither case has an appeal been sought.
Re: California governor signs AI transparency bill into law
#217Earlier quoted context omitted.
Exclusively training on 15 year source code would make code generation significantly less useful as API’s change. Economic viability and utility for AI training are closely linked. Exclude all written works including news articles etc from the last 25 years and your model will know nothing about Facebook etc. It’s not as bad if you can exclude stuff from copyright and then use that, but your proposal would have obvio…
You wouldn't need to exclusively train on 15 year old source code. What I said would simply grant you free access to all 15 year old source code, but you can already train on public domain code and likely any FOSS code without any issue, or if courts do start deciding that models inherit copyright, at the most you might have to link a list of all of the codebases you trained on with license info. The nature of the th…
Making a copy of a news article etc to train with is on the face of it copyright infringement even before you start training. Doing that for OSS is on the other hand fine, but there’s not that much OSS.
I think training itself could reasonably be considered fair use on a case by case basis. Train a neural network to just directly reproduce a work being obviously problematic etc. There’s plenty of ambiguity here.
Re: California governor signs AI transparency bill into law
#218Earlier quoted context omitted.
A judge yes, but that’s subject to appeal. The point is it never reached that stage and never will.
As an attorney, I'm trying to understand what you're getting at. His opinion, while interlocutory and not binding precedent, will be cited in future cases. And his wasn't the only one. In Kadrey v. Meta Platforms, Inc. , No. 23-cv-03417 (N.D. Cal. June 25, 2025) Judge Chhabria reached the same conclusion. https://storage.courtlistener.com/recap/gov.uscourts.cand.41... In neither case has an appeal been sought.
If you’ve read Kadrey the judge says harm from competing with the output of authors would be problematic. Quite relevant for software developers suing about code generation but much harder for novelists to prove. However, the judge came to the opposite conclusion about using pirate websites to download the books in question.
A new AI company that is expecting to face a large number of lawsuits and win some while losing others isn’t in a great position.
Re: California governor signs AI transparency bill into law
#219Earlier quoted context omitted.
Who is "we"? A lot of people don't care and will trust new AI generated stuff regardless of how wrong it is. What this means for human progress is uncertain.
And if they outcompete you, then what? I guess it'll be a little less uncertain at that point.
Re: California governor signs AI transparency bill into law
#220Earlier quoted context omitted.
I still don't really get how compensation is supposed to work just based on the math. Models are trained on billions of works and have a lifetime of around a year; AI companies (e.g. Anthropic) have revenue in the low billions of dollars a year. Even if you took all of that -- leave nothing for salaries, hardware, utilities, to say nothing of profit -- and applied it to the works in the training data, it would be app…
I think you’re overestimating the number of authors, and forgetting there’s several AI companies. A revenue sharing agreement with 10% going to creators isn’t unrealistic. Google’s revenue was 300 billion with 100 billion in profits last year, the AI industry may never reach that size but 1$/person on the planet is only 8 billion dollars, drop that to 70% of people are online so your down to 5.6 billion. That’s assum…
Google is a huge conglomerate and a poor choice for making estimates because the bulk of their revenue comes from "advertising" with no obvious way to distinguish what proportion of that ad revenue is attributable to AI, e.g. what proportion of search ad revenue is attributable to being the same company that runs the ad network, and to being the default search in Android, iOS and Chrome? Nowhere near all of it or even most of it is from AI.
"Counting books and individual Facebook posts in any language equally" is kind of the issue. The links from the AI summary things are disproportionately not to the New York Times, they're more often to Reddit and YouTube and community forums on the site of the company whose product you're asking about and Stack Overflow and Wikipedia and random personal blogs and so on.
Whereas you might have written an entire book, and that book is very useful and valuable to human readers who want to know about its subject matter, but unless that subject matter is something the general population frequently wants to know about, its value in this context is less than some random Facebook post that provides the answer to a question a lot of people have.
And then the only way anybody is getting a significant amount of money is if it's plundering the little guy. Large incumbent media companies with lawyers get a disproportionate take because they're usurping the share of YouTube creators and Substack authors and forum posters who provided more in aggregate value but get squat. And I don't see any legitimacy in having it be Comcast and the Murdoch family who take the little guy's share at the cost of significant overhead and making it harder for smaller AI companies to compete with the bigger ones.