My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…
I have a theory about why it's so easy to underestimate long-term progress and overestimate short-term progress. Before a technology hits a threshold of "becoming useful", it may have a long history of progress behind it. But that progress is only visible and felt to researchers. In practical terms, there is no progress being made as long as the thing is going from not-useful to still not-useful. So then it goes from…
Got 3.5 felt like things were improving super super fast and created that feeling the near feature will be unbelievable.
Got to 4/o series, it felt things had improved but users weren't as thrilled as with the leap to 3.5
You can call that bias, but clearly version 5 improvements displays an even greater slow down, that's 2 long years since gp4.
For context:
- gpt 3 got out in 2020
- gpt 3.5 in 2022
- gpt 4 in 2023
- gpt 4o and clique, 2024
After 3.5 things slowed down, in term of impact at least. Larger context window, multi-modality, mixture of experts, and more efficienc: all great, significant features, but all pale compared to the impact made by RLHF already 4 years ago.