Instrumentation checklist for running large GPU clusters
1–4 of 4 posts
Re: Instrumentation checklist for running large GPU clusters
#2[deleted]
Re: Instrumentation checklist for running large GPU clusters
#3Just stumbled upon this blog which details out testing and validating large GPU clusters before running training workloads. Any other similar blogs which adds more nuance in terms of debugging the issues once we identify them as well?
Re: Instrumentation checklist for running large GPU clusters
#4Just stumbled upon this blog which details out testing and validating large GPU clusters before running training workloads. Any other similar blogs which adds more nuance in terms of debugging the issues once we identify them as well?
this is a good one for debugging rdma: https://docs.redhat.com/en/documentation/red_hat_enterprise_...