Earlier quoted context omitted.
How do you figure goal generation and supervised goal training are interchangeable?
Layman warning! But "at sufficient scale", like with learning-to-learn, I'd expect it to pick up largely meta-patterns along with (if not rather than) behavioral habits, especially if the goal is left open, because strategies generalize across goals and thus get reinforcement from every instance of goal pursuit during base training. But also my intuition is that humans are "trained on goals" and then reverse-engineer…
Noticing this, frameworks like SMART[1], provide explicit generation rules. The existence of explicit frameworks is evidence that humans tend to perform worse than expected at extracting implicit structure from goals they've observed.
1. Independent of the effectiveness of such frameworks