Fully Sharded Data Parallel: Faster AI Training with Fewer GPUs
engineering.fb.com
Fully Sharded Data Parallel: Faster AI Training with Fewer GPUs
1–3 of 3 posts
Re: Fully Sharded Data Parallel: Faster AI Training with Fewer GPUs
#2So, ZeRO-Offload?
Re: Fully Sharded Data Parallel: Faster AI Training with Fewer GPUs
#3So, ZeRO-Offload?
ZeRO-Offload and ZeRO-3 and GPipe.