As an academic researcher, I find analysis of GPT-4 itself — as it pertains to other fields — to be essentially meaningless. There are no guarantees that the version of the model that was used will be available in the future (the API endpoints seem to have a ~12 month future-looking guarantee at most). Don't get me wrong: 1. GPT-4 is incredibly interesting 2. Studying GPT-4 is interesting for people working in that f…
It shows current state and progress. No strong reason to believe future models will preform worse.
But for people who are nominally using this to conduct scientific inquiries in other domains, the specific performance characteristics are what actually matter. When I am writing about the results of my semantic segmentation model, the characteristics of that specific model are more important than the notion that future models will be at least as good.
Hence my critique being pretty narrow (the academic use of GPT-4 for downstream science).