Monitor and Optimize your large-scale model training
1–3 of 3 posts
Re: Monitor and Optimize your large-scale model training
#2This is really cool! When we were trying to launch the GSPMD feature for PyTorch/XLA at Google, one of our biggest bottlenecks was network overhead, but we didn't really have any robust tools to dig into it and perform root cause analysis. I'm loving the tools I see come out of Trainy.
Re: Monitor and Optimize your large-scale model training
#3This is really cool! When we were trying to launch the GSPMD feature for PyTorch/XLA at Google, one of our biggest bottlenecks was network overhead, but we didn't really have any robust tools to dig into it and perform root cause analysis. I'm loving the tools I see come out of Trainy.
Thanks! Let me know if there are any features you'd like to see added.