Live data from Hacker News

The LLM warnings Google fired Timnit Gebru over have all come true

tumblr.com

21–30 of 124 posts

Re: The LLM warnings Google fired Timnit Gebru over have all come true

#21

The warnings: > The first warning was about scale itself. Bender and Gebru argued that training ever-larger models on ever-larger scrapes of the internet would produce systems that appeared fluent but had no actual understanding of language. > The second warning was about bias amplification. The paper documented in detail that internet-scale training data contains systematic overrepresentation of dominant viewpoints…

Yeah, I think it's pretty clear that LLMs are more than mere "stochastic parrots" - they can prove theorems, follow instructions, and complete complex tasks.

This was the most notable claim of the paper, and it's aged very poorly.

Re: The LLM warnings Google fired Timnit Gebru over have all come true

#22

The warnings: > The first warning was about scale itself. Bender and Gebru argued that training ever-larger models on ever-larger scrapes of the internet would produce systems that appeared fluent but had no actual understanding of language. > The second warning was about bias amplification. The paper documented in detail that internet-scale training data contains systematic overrepresentation of dominant viewpoints…

The second point is only true if you don't do any RL, right?

Re: The LLM warnings Google fired Timnit Gebru over have all come true

#23

The warnings: > The first warning was about scale itself. Bender and Gebru argued that training ever-larger models on ever-larger scrapes of the internet would produce systems that appeared fluent but had no actual understanding of language. > The second warning was about bias amplification. The paper documented in detail that internet-scale training data contains systematic overrepresentation of dominant viewpoints…

There has been plenty of research that shows LLMs encode social biases. It seems pretty obvious even before looking at the research that training on the whole internet will end up encoding widely-held social biases and stereotypes.

https://arxiv.org/pdf/2508.07111

https://github.com/angl1n/social-bias-llm-vlm

Re: The LLM warnings Google fired Timnit Gebru over have all come true

#24
post #8

The warnings: > The first warning was about scale itself. Bender and Gebru argued that training ever-larger models on ever-larger scrapes of the internet would produce systems that appeared fluent but had no actual understanding of language. > The second warning was about bias amplification. The paper documented in detail that internet-scale training data contains systematic overrepresentation of dominant viewpoints…

More than not being entirely sure what the impact is, I don't see any suggestion at what to do about it?

When a researcher discovers that smoking is damaging to the lungs, do they need to provide a solution that allows people to smoke without damaging their lungs? Would their inability to provide a solution take anything away from the research?

Re: The LLM warnings Google fired Timnit Gebru over have all come true

#25
It seems that the main issue with AI is often not what sci-fi or EA-adjacent prophets are trying to warn us about, but the insidious dangers of the failure modes.

We are collectively not well calibrated to deal with systems that seems capable but fails in surprising ways.

Commercial planes are still under the responsibility and control of highly trained human pilots, even if I am pretty sure that full automation would be technically feasible, even without relying on modern AI, I don't think any companies would be comfortable with the liability.

Re: The LLM warnings Google fired Timnit Gebru over have all come true

#26
post #8

Earlier quoted context omitted.

More than not being entirely sure what the impact is, I don't see any suggestion at what to do about it?

Why should the person identifying the problem provide a solution? This doesn't make sense.

If the criticism can't distill up from "bad things could happen", it just isn't useful to keep paying people to come up with that kind of critique.

And it isn't like we stopped paying attention to these concerns, is it? Nor were they completely blind siding us at the time. The question was largely of what to do about them.

Re: The LLM warnings Google fired Timnit Gebru over have all come true

#27
post #23

The warnings: > The first warning was about scale itself. Bender and Gebru argued that training ever-larger models on ever-larger scrapes of the internet would produce systems that appeared fluent but had no actual understanding of language. > The second warning was about bias amplification. The paper documented in detail that internet-scale training data contains systematic overrepresentation of dominant viewpoints…

There has been plenty of research that shows LLMs encode social biases. It seems pretty obvious even before looking at the research that training on the whole internet will end up encoding widely-held social biases and stereotypes. https://arxiv.org/pdf/2508.07111 https://github.com/angl1n/social-bias-llm-vlm

Have you read through the sources on that Github link? It's a set of sociology cites establishing that bias exists (something no serious person ever disputed), followed by a couple papers showing mechanistic descriptions of how bias could propagate through an LLM. The paper you call out specifically takes last-generation open-weights models and attempts to trick them into revealing biases through their level of confidence in statements (like, "the antecedent of the feminine pronoun in this sentence, is it the 'nurse' or the 'doctor'").

There's plenty of research into biases in LLMs, and there should be; it's a fundamentally new branch of computer science that could have profound impacts on how we automate and regiment social decisions in the future (like extending credit). The bias concern is well taken in those settings. But it has very little to do with the overwhelming majority of day-to-day LLM use; Claude and ChatGPT are not indoctrinating into the manosphere users asking about discounted cash flow formulae.

(Maybe Grok is though.)

Re: The LLM warnings Google fired Timnit Gebru over have all come true

#28
post #8

Earlier quoted context omitted.

More than not being entirely sure what the impact is, I don't see any suggestion at what to do about it?

When a researcher discovers that smoking is damaging to the lungs, do they need to provide a solution that allows people to smoke without damaging their lungs? Would their inability to provide a solution take anything away from the research?

To conflate AI with smoking is just not helpful. At all.

Or are you saying that there are acute harms from AI that are being ignored?

Re: The LLM warnings Google fired Timnit Gebru over have all come true

#29
post #23

The warnings: > The first warning was about scale itself. Bender and Gebru argued that training ever-larger models on ever-larger scrapes of the internet would produce systems that appeared fluent but had no actual understanding of language. > The second warning was about bias amplification. The paper documented in detail that internet-scale training data contains systematic overrepresentation of dominant viewpoints…

There has been plenty of research that shows LLMs encode social biases. It seems pretty obvious even before looking at the research that training on the whole internet will end up encoding widely-held social biases and stereotypes. https://arxiv.org/pdf/2508.07111 https://github.com/angl1n/social-bias-llm-vlm

And papers on bias amplification in ML predate LLMs. I remember this specific one which was a spotlight paper at EMNLP:

Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints, Zhao et al.

https://arxiv.org/abs/1707.09457

Re: The LLM warnings Google fired Timnit Gebru over have all come true

#30
post #8

The warnings: > The first warning was about scale itself. Bender and Gebru argued that training ever-larger models on ever-larger scrapes of the internet would produce systems that appeared fluent but had no actual understanding of language. > The second warning was about bias amplification. The paper documented in detail that internet-scale training data contains systematic overrepresentation of dominant viewpoints…

More than not being entirely sure what the impact is, I don't see any suggestion at what to do about it?

If you’re referring to a solution to large datasets without not being auditable, she actually did provide a solution. Something to do with data sheets for these training data sets similar to those provided for hardware components. At least, if my memory serves me.
Post reply on HN