Live data from Hacker News

California governor signs AI transparency bill into law

gov.ca.gov

171–180 of 232 posts

Re: California governor signs AI transparency bill into law

#171

Earlier quoted context omitted.

Protection for whistleblowers - which might expose nefarious actions

I think protection for whistleblowers both in AI and in general is a good thing, but ... do we really need a special carveout for AI whistleblowers? Do we not already have protections for them, or is it insufficient? And if we don't have them already, why not pass general protections instead of something so hyper-specific? (not directing these questions at you specifically, though if you know I'd certainly love to he…

You have to understand what the purpose of this bill is.

It's not supposed to do anything in particular. It's supposed to demonstrate to the public that lawmakers are Taking Action about this whole AI thing.

An earlier version of the bill had a bunch of aggressive requirements, most of which would have been bad. The version that passed is more along the lines of filing paperwork and new rules that are largely redundant with existing rules, which is wasteful and effectively useless. But that was the thing that satisfied the major stakeholders, because the huge corporations don't care about spending ~0% of their revenue on some extra paper pushers and the legislators now get to claim that they did something about the thing everybody is talking about.

Re: California governor signs AI transparency bill into law

#172

Earlier quoted context omitted.

The problem it solves is providing any sort of baseline framework for lawmakers and the legal system to even discuss AI and its impacts based on actual data instead of feels. That's why so much of it is about requiring tech companies to publish safety plans, transparency reports and incidents, and why the penalty for noncompliance is only $10,000. A comprehensive AI regulatory action is way too premature at this stag…

> A comprehensive AI regulatory action is way too premature at this stage Funny, I think it is overdue.

> I think it is overdue.

Why?

Re: California governor signs AI transparency bill into law

#173

Earlier quoted context omitted.

What else I could possibly do if complying is not technically possible?

well certainly not geoblock a lot of your customers maybe your situation is different, but if we geoblocked all of california we'd go out of business within a year

For global companies, California is less than 40 million people out of more than 8 billion. For regionally concentrated companies, they only have to care if the region they're concentrated in is California. For everyone else, losing a fraction of a percent of the customer base in exchange for lower compliance costs and legal risks is often completely logical.

Re: California governor signs AI transparency bill into law

#174
post #167

Earlier quoted context omitted.

A bubble's when values inflate well beyond intrinsic valuations. You can have a crash without this having happened.

Yes, indeed. Thought I'd change "intrinsic valuations" to "rapid change in valuations". If you deflate over time slowly (or simply slow growth expectations) you can prevent a pop. That doesn't tend to be what historically happens, though.

> to "rapid change in valuations".

Yes. This seems to be what you're doing-- carrying the metaphor too far. I believe the common use relates the asset prices to intrinsic or fundamental values.

Re: California governor signs AI transparency bill into law

#175

I found this website with the actual bill text along with annotations [0]. The section 22757.12. seems to contain the actual details of what they mean by "transparency". [0] https://sb53.info/

> “Artificial intelligence model” means an engineered or machine-based system that varies in its level of autonomy and that can, for explicit or implicit objectives, infer from the input it receives how to generate outputs that can influence physical or virtual environments. Correct me if I'm wrong, but it sounds like this definition covers basically all automation of any kind. Like, a dumb lawnmower responds to the…

Motion sensing light? Or light sensing? Motion sensors say with automated doors...

They infer from fuzzy input when to activate...

Re: California governor signs AI transparency bill into law

#176
post #130

Earlier quoted context omitted.

No, but society has no reason to grant monopolies on 50 year old publications (e.g. textbooks or news articles written or songs recorded prior to 1975), and the changes that were made to copyright law to extend it into multiple generations were actual rent-seeking. i.e. manipulating public policy to transfer wealth from others to you rather than creating wealth. Going with the original 28 year compromise from a time…

It’s far more recent works that AI companies care about. They can’t copy 50 year old python, JavaScript, etc code because it simply doesn’t exist. There’s some 50 year old C code, but it’s no longer idiomatic and so it goes. Utility of older works drop off as science marches on and culture changes. The real secret of long copyright terms is they just don’t matter much. Steamboat Willy entered the public domain and fo…

I still don't really get how compensation is supposed to work just based on the math. Models are trained on billions of works and have a lifetime of around a year; AI companies (e.g. Anthropic) have revenue in the low billions of dollars a year.

Even if you took all of that -- leave nothing for salaries, hardware, utilities, to say nothing of profit -- and applied it to the works in the training data, it would be approximately $1 each.

What is that good for? It would have a massive administrative cost and the authors would still get effectively nothing.

Re: California governor signs AI transparency bill into law

#177
post #94

Earlier quoted context omitted.

> Add in the narrow exceptions like child porn and true threats, and that's it. You're contradicting yourself. On the one hand you're saying that governments shouldn't have the power to define "safety", but you're in favor of having protections against "true threats". How do you define "true threats"? Whatever definition you may have, surely something like it can be codified into law. The questions then are: how loos…

There's no contradiction. "True threats" is already a narrow exception defined by decades of Supreme Court precedent. It means statements where the speaker intends to communicate a serious expression of intent to commit unlawful violence against a person or group. That's it. It's not a blank check for the government to decide what counts as dangerous. Brandenburg gives us the standard: speech can only be restricted i…

I'm not sure why you're only focusing on speech. "True threats" doesn't come close to covering all the possible use cases and ways that "AI" tools can be harmful to society. We can't apply legal precedent to a technology without precedent.

> "This technology is different" is what every regulator says about every new technology. Print was different. Radio was different. The internet was different.

"AI" really is different, though. Not even the internet, or computers, for that matter, had the potential to transform literally every facet of our lives. Now, I personally don't buy into the "AGI" nonsense that these companies are selling, but it is undeniable that even the current generation of these tools can shake up the pillars of our society, and raise some difficult questions about humanity.

In many ways, we're not ready for it, yet the companies keep producing it, and we're now deep in a global arms race we haven't experienced in decades.

> I want the government to stay out of mandating content restrictions. Not because I trust corporations, but because I trust the government even less with the power to define what information is too dangerous to share.

See, this is where we disagree.

I don't trust either of them. I'm well aware of the slippery slope that is giving governments more power.

But there are two paths here: either we allow companies to continue advancing this technology with little to no oversight, or we allow our governments to enact regulation that at least has the potential to protect us from companies.

Governments at the very least have the responsibility to protect and serve their citizens. Whether this is done in practice, and how well, is obviously highly debatable, and we can be cynical about it all day. On the other hand, companies are profit-seeking organizations that only serve their shareholders, and have no obligation to protect the public. In fact, it is pretty much guaranteed that without regulation, companies will choose profits over safety every time. We have seen this throughout history.

So to me it's clear that I should trust my government over companies. I do this everyday when I go to the grocery store without worrying about food poisoning, or walk over a bridge without worrying that it will collapse. Shit does happen, and governments can be corrupted, but there are general safety regulations we take for granted every day. Why should tech companies be exempt from it?

Modern technology is a complex beast that governments are not prepared to regulate. There is no direct association between technology and how harmful it can be; we haven't established that yet. Even when there is such a connection, such as smoking causing cancer, we've seen how evil companies can be in refuting it and doing anything in their power to preserve their revenues at the expense of the public. "AI" further complicates this in ways we've never seen before. So there's a long and shaky road ahead of us where we'll have to figure out what the true impact of technology is, and the best ways to mitigate it, without sacrificing our freedoms. It's going to involve government overreach, public pushback, and company lobbying, but I hope that at some point in the near future we're able to find a balance that we're relatively and collectively happy with, for the sake of our future.

Re: California governor signs AI transparency bill into law

#178

I found this website with the actual bill text along with annotations [0]. The section 22757.12. seems to contain the actual details of what they mean by "transparency". [0] https://sb53.info/

> “Artificial intelligence model” means an engineered or machine-based system that varies in its level of autonomy and that can, for explicit or implicit objectives, infer from the input it receives how to generate outputs that can influence physical or virtual environments. Correct me if I'm wrong, but it sounds like this definition covers basically all automation of any kind. Like, a dumb lawnmower responds to the…

"The Butlerian Jihad was a cataclysmic, millennia-long holy war that completely eradicated artificial intelligence, computers, and sentient robots from human civilization. Taking place over 10,000 years before the events of the original novel, the Jihad was a violent reaction to humanity's over-reliance on and eventual enslavement by "thinking machines". The ban on advanced AI became a foundational law of the Galactic Imperium."

Re: California governor signs AI transparency bill into law

#179

I found this website with the actual bill text along with annotations [0]. The section 22757.12. seems to contain the actual details of what they mean by "transparency". [0] https://sb53.info/

> “Artificial intelligence model” means an engineered or machine-based system that varies in its level of autonomy and that can, for explicit or implicit objectives, infer from the input it receives how to generate outputs that can influence physical or virtual environments. Correct me if I'm wrong, but it sounds like this definition covers basically all automation of any kind. Like, a dumb lawnmower responds to the…

Yeah, you're wrong - a court simply isn't going to consider a lawnmower's translation of throttle input to motor power as "inference". The principles of statutory interpretation require courts to consider the context and purpose of the legislation, and everyone knows this is about GPT-5, not lawnmowers.

In any case, that definition is only used to further define "foundation model": "an artificial intelligence model that is all of the following: (1) Trained on a broad data set. (2) Designed for generality of output. (3) Adaptable to a wide range of distinctive tasks." This legislation is very clearly not supposed to cover your average ML classifier.

Re: California governor signs AI transparency bill into law

#180
post #36

Earlier quoted context omitted.

I'd rather not pay OpenAI either. I'll stick with my open-weights models, and I rather anachronistic rent-seeking not kill those. You're not getting a cent from OpenAI, and the government isn't going to do anything about it. Just get over it.

All of them are trained on copyrighted data. Why is it okay for a model to serve up paraphrased books but verbatim copies from the Pirate Bay are illegal? I don’t deny the utility of LLMs. But copyright law was meant to protect authors from this kind of exploitation. Imagine instead of “magical AGI knowledge compression”, instead these LLM providers just did a search over their “borrowed” corpus and then performed a…

> Why is it okay for a model to serve up paraphrased books but verbatim copies from the Pirate Bay are illegal?

Because they are not actually memorizing those books (besides few isolated pathological cases due to imperfect training data deduplication), and whatever they spit out is in no way a replacement for the original?

Here's some back-of-the-envelope math: Harry Potter and the Philosopher's Stone is around ~460KB of text and equivalent to ~110k Qwen3 tokens, which gives us ~0.24 tokens per byte. Qwen3 models were trained on 36 trillion tokens, so this gives us a dataset of ~137TB. The biggest Qwen3 model has ~235B parameters and at 8-bit (at which you can serve the model essentially loselessly compared to full bf16 weights) takes ~255GB of space, so the model is only 0.18% of its training dataset. And this is the best case, because we took the biggest model, and the actual capacity of a model to memorize is only at most ~4 bit per parameter[1] instead of full ~8 bits we assumed here.

For reference, the best loseless compression we can achieve for text is around ~15% of the original size (e.g. Fabrice Bellard's NNCP), which is two orders of magnitude worse.

So purely from information theoretic perspective saying that those models memorize the whole datasets on which they were trained is nonsense. They can't do that, because there's just not enough bits to store all of this data. They extract patterns, the same way that I can take the very same ~137TB dataset and build a frequency table of all of the bigrams appearing in it and build a hidden Markov model out of it to generate text. Would that also be "stealing"? And what if I extend my frequency table to trigrams? Where exactly do we draw the line?

[1] -- https://arxiv.org/pdf/2505.24832

Post reply on HN