Eleuther.ai is just a bunch of random, but smart people without capital who decided on Twitter to recreate GPT-3. Recently they released GPT-NeoX-20B. They mainly coordinate on Discord. They got compute from some company for free. https://www.eleuther.ai/ Another group called BigScience got a grant from France to use a public institution supercomputer to train large language model in open. They are 71% done training…
I don't trust papers out of “Top Labs” anymore
21–30 of 53 posts
Re: I don't trust papers out of “Top Labs” anymore
#22NNs, and models of this kind, are just search engines. They store a compression of of everything ever written, and prediction is just googling through it.
Models performance exponential in parameter count should be just ignored by research. This category of performance is already established by research, more compute and more historical data stored, isnt an interesting research result.
Re: I don't trust papers out of “Top Labs” anymore
#23At least for ML you can mostly reproduce the results, even in if they’re not that interesting.
Re: I don't trust papers out of “Top Labs” anymore
#24Re: I don't trust papers out of “Top Labs” anymore
#25They explicitly say they trust the results. They're complaining that top labs use lots of compute, so the results aren't relevant to someone who can't. They give an example where a paper used 18K TPU core hours. It's easy to find papers that use millions of core hours. IMO, asking AI people to not use expensive compute is like asking astronomers to please stop using expensive telescopes. The opposite side of this arg…
Re: I don't trust papers out of “Top Labs” anymore
#26Re: I don't trust papers out of “Top Labs” anymore
#27At this tiny number, the randomness is starting to affect the scores. Like labeling mistake of test data by human. Maybe, training SotA with different random seeds make its score 0.03% better or worse.
Hell, 17,810 TPU core-hours is a huge number. You can't ignore the work of randomness. What if a cosmic ray hit a specific memory cell which cause the soft memory error, causing a single wrong calculation which ultimately cause the final trained model 0.03% difference?
So, it's more like: "Jeff Dean spent enough money to feed a family of four for half a decade to get a 0.03% of winning lottery on CIFAR-10."
Re: I don't trust papers out of “Top Labs” anymore
#28They explicitly say they trust the results. They're complaining that top labs use lots of compute, so the results aren't relevant to someone who can't. They give an example where a paper used 18K TPU core hours. It's easy to find papers that use millions of core hours. IMO, asking AI people to not use expensive compute is like asking astronomers to please stop using expensive telescopes. The opposite side of this arg…
We already know that more compute hours give better results, and a paper which simply consists of rerunning previous work but with 100x the compute hours for .03% better results has not discovered anything new, and there's no point in reading that paper.
Re: I don't trust papers out of “Top Labs” anymore
#29They explicitly say they trust the results. They're complaining that top labs use lots of compute, so the results aren't relevant to someone who can't. They give an example where a paper used 18K TPU core hours. It's easy to find papers that use millions of core hours. IMO, asking AI people to not use expensive compute is like asking astronomers to please stop using expensive telescopes. The opposite side of this arg…
We already know that more compute hours give better results, and a paper which simply consists of rerunning previous work but with 100x the compute hours for .03% better results has not discovered anything new, and there's no point in reading that paper.
Re: I don't trust papers out of “Top Labs” anymore
#30They explicitly say they trust the results. They're complaining that top labs use lots of compute, so the results aren't relevant to someone who can't. They give an example where a paper used 18K TPU core hours. It's easy to find papers that use millions of core hours. IMO, asking AI people to not use expensive compute is like asking astronomers to please stop using expensive telescopes. The opposite side of this arg…
We already know that more compute hours give better results, and a paper which simply consists of rerunning previous work but with 100x the compute hours for .03% better results has not discovered anything new, and there's no point in reading that paper.