The improvement over ChatGPT are counted in (very) few percents. Does it mean they have entered a diminishing returns phase or is it that each percent is much harder to get compared to the previous ones ?
I doubt LLMs are close to plateauing in terms of performance unless there's already an awful lot more to GPT-4's training than is understood. It seems like even simple stuff like planning ahead (e.g. to fix "hallucinations", aka bullshitting) is still to come.