Earlier quoted context omitted.
It wouldn't even have to be a language model, just some sort of meta program that handled the switchboarding, as it were, between subsystems would be enough. But my real point is just more narrow, you can't measure progress to artificial general intelligence with a list of tasks satisfied by artificial specific intelligence. You have to measure the generalness directly. And thats also why I don't think you'd have to…
It’s very difficult to come up with an objective metric or benchmark for general AI that can’t be gamed, or which won’t turn out to be disappointingly easy. It makes sense that most research would be in the direction of tasks which are easily quantifiable. More hazy benchmarks like the Turing test are possibly better but that one in particular isn’t so good unless it’s enhanced (I’d say a 1 hour conversation with an…
2. You could also rework this list and say that an AI which can fulfill three tasks even badly is one step forwards. An AI which can fulfill three tasks from three separate categories is another step, etc. But I think it's a categorical mistake to count progress in specific AI as progress in general AI, there's not a very good reason to believe they are in a continuum with one another.