Earlier quoted context omitted.
>performance on standardized tests? That doesn't necessarily seem like the best metric for what the LLM tries to be. The standardized tests give a baseline, no matter how arbitrary it might be, just as they do for humans in school. Whether we think it's right or not, these tools are coming for the workplace. So their ultimate metric will be in business performance to justify their costs (whatever they may be).
GPT 3.5 had trouble understanding when I told it "Say 2 bob are a beb, how many beb per bob are there?" and it wrote a goddamn essay about shoes. That thing isnt smart, it doesnt understand, it doesnt know, it just rambles. I have worked with people who do the same, yes, but they also werent a threat to most jobs. I said it before, and I will say it again: If ChatGPT 3,4,5,... can take your job, maybe youre not reall…
Did you use the quoted prompt exactly?