Live data from Hacker News

I don't trust papers out of “Top Labs” anymore

old.reddit.com

11–20 of 53 posts

Re: I don't trust papers out of “Top Labs” anymore

#12
post #4
post #3

Quoted post unavailable.

I actually find that smaller communities on Reddit are pretty good. Just unsub from all the defaults and find your place. Talking about Reddit as a whole is not easy. It’s insane how wide it is.

The small and good ones all go to shit eventually. Eventually the karma-addicts figure out how to pander for points and turn the subreddit into a flanderization of itself, using a torrent of low-effort nominally on-topic posts intentioned only to be upvoted. The fundamental dynamics of reddit result in a 'game' where karma addicts use low-effort content to actively sabotage communities. The only way to avoid this is to keep a subreddit small and obscure, balanced on a knife edge between too-obscure to survive and popular enough to attract karma farmers.

A good and small subreddit exists in an unstable equilibrium on borrowed time.

Re: I don't trust papers out of “Top Labs” anymore

#14
Eleuther.ai is just a bunch of random, but smart people without capital who decided on Twitter to recreate GPT-3.

Recently they released GPT-NeoX-20B. They mainly coordinate on Discord. They got compute from some company for free.

https://www.eleuther.ai/

Another group called BigScience got a grant from France to use a public institution supercomputer to train large language model in open. They are 71% done training their 176 billion parameters open-source language model called "BLOOM".

> During one-year, from May 2021 to May 2022, 900 researchers from 60 countries and more than 250 institutions are creating together a very large multilingual neural network language model and a very large multilingual text dataset on the 28 petaflops Jean Zay (IDRIS) supercomputer located near Paris, France.

https://bigscience.huggingface.co/

If there is a will there is a way.

BTW - People close to EleutherAI are looking for people wanting to play around with open-source machine learning for biology.

You just need to start contributing on their Discord: https://twitter.com/nc_znc/status/1530545001557643265

Re: I don't trust papers out of “Top Labs” anymore

#15
post #2

"Jeff Dean spent enough money to feed a family of four for half a decade to get a 0.03% improvement on CIFAR-10." Nailed it.

That quote is disingenuous. Do people really think that...

* Jeff Dean, lead of Google's AI division, wrote a paper with all that complexity to get SOTA on CIFAR-10?

* Jeff Dean, whose salary is sometimes estimated as $3m/y and is responsible for the direction of research of many more, is unreasonable for using * going from a 0.6% error rate to a 0.57% error rate is reasonably summarized as ‘a 0.03% improvement’, ignoring both that it's a 5% reduction in error and that such improvements get harder as you approach (or exceed) the label accuracy of the dataset?

* the accuracy from this paper came purely from scale?

Re: I don't trust papers out of “Top Labs” anymore

#16
post #2

"Jeff Dean spent enough money to feed a family of four for half a decade to get a 0.03% improvement on CIFAR-10." Nailed it.

Put another way though, the failure rate was decreased from 0.6 to 0.57 or a 5% reduction. That's pretty significant. If you can reduce LASIK failure rate by 5%, that would provide a ton of value although you would be talking about an absolute improvement of 0.001% in success rate.

I agree that the improvements we are seeing are increasingly due to simply spending more time/money/power but that quip is probably the weakest argument. I would have liked to have seen a Fermi calculation where the power used during training is only 1% (or probably much less) of the total power used. The other thing that reeks naivety is basically the world takes a lot of compute. Much more money and compute is wasted on Candy Crush for instance.

Re: I don't trust papers out of “Top Labs” anymore

#18
post #6

It does seem like better algorithms to get similar results from smaller models should be prioritised. Rather than throwing more compute at a problem for 0.03 better score, show me one tenth the compute with a loss of 0.03 score. That would be impressive and far more useful.

While I am inclined to personally agree with your sentiment, I don't think I have better insights than Richard Sutton: http://incompleteideas.net/IncIdeas/BitterLesson.html

"The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin."

Re: I don't trust papers out of “Top Labs” anymore

#19
They explicitly say they trust the results. They're complaining that top labs use lots of compute, so the results aren't relevant to someone who can't. They give an example where a paper used 18K TPU core hours. It's easy to find papers that use millions of core hours.

IMO, asking AI people to not use expensive compute is like asking astronomers to please stop using expensive telescopes. The opposite side of this argument is "Gee, it looks like increasing compute helps AI a lot. Why the heck have we been spending so little on compute?" [0].

[0]: https://www.gwern.net/Scaling-hypothesis

Post reply on HN