Earlier quoted context omitted.
Did you even read through the issue? I don't see anything that indicates it won't work.
Yes, I did. The last comment is a traceback and an explanation what would have to be done to fix it.
Zero-3 Offload: Scale DL models to trillion parameters without code changes
31–40 of 49 posts
Re: Zero-3 Offload: Scale DL models to trillion parameters without code changes
#32Earlier quoted context omitted.
You ideally need ~500GB of text, or so. EleutherAI's The Pile was designed to be just big enough to fit a 1t GPT efficiently, and you can get the various scaling curves out of the OA-related scaling papers. (You want the amount of data that fits into a single epoch, because if you reuse data, you get less bang for the FLOPs buck, and FLOPS constraints are right now much more binding than data or model size.)
This feels off by a couple of orders of magnitude, unless a significant number of the parameters are not independent.
Re: Zero-3 Offload: Scale DL models to trillion parameters without code changes
#33Earlier quoted context omitted.
You ideally need ~500GB of text, or so. EleutherAI's The Pile was designed to be just big enough to fit a 1t GPT efficiently, and you can get the various scaling curves out of the OA-related scaling papers. (You want the amount of data that fits into a single epoch, because if you reuse data, you get less bang for the FLOPs buck, and FLOPS constraints are right now much more binding than data or model size.)
This feels off by a couple of orders of magnitude, unless a significant number of the parameters are not independent.
Re: Zero-3 Offload: Scale DL models to trillion parameters without code changes
#34This is also being added to pytorch https://github.com/pytorch/pytorch/pull/46750
Re: Zero-3 Offload: Scale DL models to trillion parameters without code changes
#35Re: Zero-3 Offload: Scale DL models to trillion parameters without code changes
#36Re: Zero-3 Offload: Scale DL models to trillion parameters without code changes
#37Alternatively, one could get rid of the memory used by optimizers entirely by switching to vanilla SGD. I haven’t tried this on transformers and maybe that’s what breaks down here but in “classic” supervised settings I’ve found SGD with schedule tuning just as fast as Adam.
SGD doesn't work on large Transformers, no. You need something like AdamW.
Re: Zero-3 Offload: Scale DL models to trillion parameters without code changes
#38Earlier quoted context omitted.
But it's not clear how they managed to improve training on a single GPU: they say they can fit 40B model on a single V100.
They offload parameters, gradients and optimizer states (such as moment, velocity and exponential avg of these in Adam) into CPU memory.
Re: Zero-3 Offload: Scale DL models to trillion parameters without code changes
#39GPT-NeoX is an example project that is using deepspeed and Zero-3 offloading. The wider project intend to train a GPT-3 sized model and release it freely to the world. https://github.com/EleutherAI/gpt-neox
It seems like Zero-3 doesn't work for them: https://github.com/EleutherAI/gpt-neox/issues/171
https://github.com/microsoft/DeepSpeed/issues/846
Also, the specific problem described in that Issue was due to a bug I found in DeepSpeed that has since been corrected.
Re: Zero-3 Offload: Scale DL models to trillion parameters without code changes
#40Earlier quoted context omitted.
Third paragraph or so in the overview: > ZeRO removes the memory redundancies across data-parallel processes by partitioning the three model states (optimizer states, gradients, and parameters) across data-parallel processes instead of replicating them. By doing this, it boosts memory efficiency compared to classic data-parallelism while retaining its computational granularity and communication efficiency
Yeah that would be the techno-babble. I've been working on a machine learning pipeline for 6 years and I still have no idea what this means.
For anyone who still thinks the blog post has substance: Saying that they partitioned the optimizer state, params, etc to have no redundancies is kind of "duh," sort of like saying "we solved the problem using coding and algorithms." It's obvious that we want to eliminate redundancies to maximize the effective VRAM; it's not like nobody thought of not having redundancies before. The problem is that in general, training models distributed is a weird balancing act between redundancy, network usage, compute, etc. The existing methods, model/pipeline/data parallel, gradient checkpoint/accum, etc all have their pros and cons. Unless ZeRO3 is doing something crazy, it has to be giving something up to get to zero redundancy, and knowing what that is would be very important.
If someone could ELI5 how ZeRO actually works, that would be nice.