Live data from Hacker News

We read the paper that forced Timnit Gebru out of Google

technologyreview.com

411–420 of 683 posts

Re: We read the paper that forced Timnit Gebru out of Google

#411

Are these numbers the energy to train a model? The whole point of these new NLP models is transfer learning, meaning you train the big model once and fine-tune it for each use case with a lot less training data. 5 cars worth of carbon emissions is not a lot given that it is a fixed cost. Very few are retraining BERT from scratch. EDIT: The other two points are also disingenuous. * "[AI models] will also fail to captu…

> 5 cars worth of carbon emissions is not a lot given that it is a fixed cost. Very few are retraining BERT from scratch. It's not a fixed cost though, it's just how much was spent on this year's iteration of the model. The overall point being made is that model training costs are growing unbounded. Next year it could be 30 cars' worth or whatever for BERT-2, then 600 cars' worth for BERT-3 the year after. That's wha…

This is a very valid argument but it's hard to know what scaling a transformer will really do without trying (looking at you GPT-3). This is probably an issue for ML in general at this point.

I think a more nuanced conversation around these topics will look at exactly what you bring up, how do we properly trade the potential knowledge benefit against the costs?

It pains me that entirely valid avenues of research like this get covered up in nonsense and drama and their message seemingly lost in the midst of it.

Re: We read the paper that forced Timnit Gebru out of Google

#412

Earlier quoted context omitted.

Toxic is an inflammatory and unnecessary word to use when no-one is privy to the actual facts.

A word you seem to have no reservations about using against your opponents just from a quick search of your comment history. Why such outrage when it's turned back on your own sacred cows?

Why are you assuming anything about me? You realise I e used the word toxic in my whole comment history only related to this topic right?

I, like many others, dont like the way she deals with people. It is toxic. You call a toxic person, a toxic person. Ample evidence for it. Its not outrage, its just facts.

Re: We read the paper that forced Timnit Gebru out of Google

#413

Are these numbers the energy to train a model? The whole point of these new NLP models is transfer learning, meaning you train the big model once and fine-tune it for each use case with a lot less training data. 5 cars worth of carbon emissions is not a lot given that it is a fixed cost. Very few are retraining BERT from scratch. EDIT: The other two points are also disingenuous. * "[AI models] will also fail to captu…

The Strubell paper which is the origin of this "5 cars" number isn't even in the right ballpark for this stuff. What they did was take desktop GPU power consumption running the model in fp32, extrapolate to a 240x GPU (P100) setup that would run for a year straight at 100% power consumption. Yes if you do run 240x p100s at literally 100% 24/7 for a year you get the power consumption of 5 cars. This run never happened…

> I've never profiled world-wide carbon production but something tells me if you wanted to carbon optimise you'd be better served trying to take cars off the road and planes out of the sky.

We're getting a bit off-topic here, but the #1 target by far in reducing greenhouse emissions is power generation. In transportation it's significantly trickier to replace petrol-based fuels (especially for airplanes), but it's straightforward enough in power plants. And crucially, you can convert all the petrol-powered vehicles to EVs that you want, but if the electricity they're getting from the wall is still provided by burning petroleum then you haven't actually done that much.

Luckily, computation can for the most part be located anywhere (exactly the opposite of transportation), and thus you have a lot of data centers near hydro and other renewable sources so that they can use the cheapest green power available.

Re: We read the paper that forced Timnit Gebru out of Google

#414

Earlier quoted context omitted.

> 5 cars worth of carbon emissions is not a lot given that it is a fixed cost. Very few are retraining BERT from scratch. It's not a fixed cost though, it's just how much was spent on this year's iteration of the model. The overall point being made is that model training costs are growing unbounded. Next year it could be 30 cars' worth or whatever for BERT-2, then 600 cars' worth for BERT-3 the year after. That's wha…

This is a very valid argument but it's hard to know what scaling a transformer will really do without trying (looking at you GPT-3). This is probably an issue for ML in general at this point. I think a more nuanced conversation around these topics will look at exactly what you bring up, how do we properly trade the potential knowledge benefit against the costs? It pains me that entirely valid avenues of research like…

Yes, an in-depth nuanced conversation is absolutely merited here, and Gebru and her colleagues were having it in the most rigorous way possible -- in peer-reviewed papers at AI conferences. I won't pretend to remotely be contributing as much to the discussion as they were.

Re: We read the paper that forced Timnit Gebru out of Google

#415
post #13
post #12

This doesn't strike me as anything that needs to be buried. The energy argument is tenuous and at this point models use relatively little compute power. The language bias is a better tack but I don't think this is earth shattering to anyone. Most of the internet is from Western sources and some fraction of that is racist/prejudiced. I think this is obvious to anyone who has ever used the internet.

It sounds like it should be buried because it's clickbait headline generating fluff that only would get attention because of where the authors come from. And I say that after reading a summary that sounds like the authors had a positive opinion of the paper. Google (probably) wanted to bury it because many of those clickbait headlines would have been negative to google.

Reserve your judgment until you can read the actual paper.

Re: We read the paper that forced Timnit Gebru out of Google

#416

Are these numbers the energy to train a model? The whole point of these new NLP models is transfer learning, meaning you train the big model once and fine-tune it for each use case with a lot less training data. 5 cars worth of carbon emissions is not a lot given that it is a fixed cost. Very few are retraining BERT from scratch. EDIT: The other two points are also disingenuous. * "[AI models] will also fail to captu…

I'm more interested in the related nugget, which is the carbon footprint of a Google Search. Some estimates from a decade back put it at 7g (boil a cup of water), but since then it's probably only gotten larger. However if Google's dstacenters truly carbon neutral, does it even matter?

Green energy is often used in addition to the carbon-based energy. The total amount of used energy increases. Green energy should replace carbon-base energy.

Wind and solar are finite too. The places where they can be harvested are scarce. So if Google using up green energy for bells and whistles on the search page, less homes can use green energy for heating and transport.

Re: We read the paper that forced Timnit Gebru out of Google

#417

Anyone else see the narrative on Twitter to be so much different than Reddit and hacker news? On reddit and hacker news I have never seen any unconditional support for Timmit. Even amongst those anti Google and broadly in support of her, there's no whole agreement with all of what Timmit says; the assertion of her story, the reasons why and the conclusions. that a. She was fired b. That she was fired because of sexis…

I find that the average HN thread demonstrates a broader diversity of thought, and generally deeper, more critical, more scientific thought than the average Twitter feed. So it’s not surprising that Twitter would be near unanimous and HN would be more nuanced. This is one of the reasons I prefer HN to just about any other online forum.

I agree that there is a broader diversity of a thought, but there’s also a dominant, conservative streak when it comes to deferring to corporate authority. This is unsurprising since many people on this site get paid a lot of money by the same people being critiqued, but does represent a kind of self-selection toward servility.

Re: We read the paper that forced Timnit Gebru out of Google

#418

Earlier quoted context omitted.

The Strubell paper which is the origin of this "5 cars" number isn't even in the right ballpark for this stuff. What they did was take desktop GPU power consumption running the model in fp32, extrapolate to a 240x GPU (P100) setup that would run for a year straight at 100% power consumption. Yes if you do run 240x p100s at literally 100% 24/7 for a year you get the power consumption of 5 cars. This run never happened…

> I've never profiled world-wide carbon production but something tells me if you wanted to carbon optimise you'd be better served trying to take cars off the road and planes out of the sky. We're getting a bit off-topic here, but the #1 target by far in reducing greenhouse emissions is power generation. In transportation it's significantly trickier to replace petrol-based fuels ( especially for airplanes), but it's s…

> We're getting a bit off-topic here, but the #1 target by far in reducing greenhouse emissions is power generation.

I readily admit I don't know any of the numbers associated with carbon production and my comment was solely based on the one GPU vs car figure presented in the aforementioned paper.

Re: We read the paper that forced Timnit Gebru out of Google

#419
post #90

Earlier quoted context omitted.

The issue isn't how much BERT uses, the issue is the trend over time in how much models use, with BERT being a recent data point. The whole point of having AI ethicists is to identify current indicators of potential future ethical problems so that they can be considered in guiding the direction of development, so that you minimize acute ethical crisis.

You do realize just how small this co2 output is compared to the human co2 output it replaces. You realize the shear scale of co2 output the office the engineers who wrote the model, drove to work, education, etc produced. I’m actually suprised just how small of a co2 output it was.

The number is probably quite a bit lower. The Strubell paper uses a PUE coefficient of 1.58, while Google datacenters are at 1.1. The figures are for GPUs, but Google uses TPUs, whose power characteristics were not public. Price is a proximate for power usage, though. TPUs might have been half as expensive in that experiment. Let's say that a further half of those gains are money that goes to Nvidia and not really related to power. So, making numbers up, training with TPUs might be another 25% more efficient. That's probably conservative, as Google claims that the TPUs are tens of times less power-hungry, due to their simplicity, but on the other hand, you also other fixed costs like racks, fabric and the CPUs to feed the chips.

In an ideal world, nobody would get hung up on details and everybody would understand that there is a lot of nuance when comparing things. In practice, if Google published a paper which quoted the Strubell paper without the caveats, I can see headlines about how inefficient and bad for the environment Google Translate is. PR would get busy and obtain corrections or follow-up articles to clarify things, but those rarely get the same attention. And it's still extra work that could have been avoided, which in itself is bad optics ("do you folks even review the stuff you send out for publication?").

I'm all for reducing emissions and improving efficiency, but I find the premise a bit of a stretch.

Yes, large and wealthy organizations have a big advantage, but that applies to pretty much anything they do, not just language models. Inefficiencies are bad, but if you ask a number of people why fix them, I think they'd mention financial cost and wasted time well before looking at it as an issue of ethics and fairness.

Reducing costs for language models is already a great idea across all fronts. So is reducing inequalities. It's linking the two that sounds like a strained argument to me. Suppose someone makes models ten times smaller next week. Will marginalized communities' lives improve soon? There must be more. I haven't seen the paper, so I'm curious how it is all framed.

Post reply on HN