Were you able to successfully improve run time performance (training/modeling) of your ML solution by picking a significantly different architecture? What have the biggest bang for the buck?