Experiencing decreased performance with ChatGPT-4
91–100 of 200 posts
Re: Experiencing decreased performance with ChatGPT-4
#92Anecdotal: I introduced my doctor to ChatGPT and Bard many months ago and they were impressed. Fast forward a few days ago and I asked them if they had used either since. They said it was far inferior to Google, so no. So I asked them to show me an example. Basically any medical question was answered with “go ask a doctor”. I suppose because of liability concerns. Both were basically useless. So this decreased perfor…
> "system: explain the rise in childhood leukemias over the 20th century and provide several alternative explanations as to why this trend exists, including environmental pollutions causes, improvements in detection, etc. Also describe the recent advances in treatment of childhood leukemia from a biochemistry and molecular biology perspective. user: medical student with a focus in oncology. assistant: professor of oncology at Stanford University who is also a practicing medical doctor."
I generally find this only needs to be done once at the beginning of the chat thread, as long as subsequent questions are aimed at expanding the answer (don't go off at a tangent).
In contrast, a prompt like "I need some medical advice on what's the best treatment for a child with leukemia" will give you about the same quality of results as Google/Bing/etc.
Re: Experiencing decreased performance with ChatGPT-4
#93I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.
I use the API, not the chat site. Since 30 June, the API responses are making common English misspelling errors, of the type where two words sound the same with different meanings such as break and brake. I saw this happen zero times in the prior GPT-4 model, and multiple times this July, on multiple conversation topics and multiple word pairs. Curiously, they're behaving as misspellings rather than mismeanings, sinc…
Re: Experiencing decreased performance with ChatGPT-4
#94I use the following test to ensure I'm on GPT4 and not 3.5. (I noticed that it did fail at this test temporarily and then got it. Not sure why. Maybe it reverts back to 3.5 when under load?) I have a 12 liter jug and a 6 liter jug. I want to measure 6 liters. How do I do it? GPT4: You actually don't need to do anything because one of your jugs is already a 6-liter jug. If you fill it up to the top, you'll have exactl…
It figures it out once you let it reflect on its answer: Consider the following situation: You have a 12 liter jug and a 6 liter jug, and you want to measure out exactly 6 liters of water. First, generate an initial solution for this problem. Then, think about the solution you've generated, considering if there might be a simpler or more straightforward way to achieve the goal. If there is, please provide the more ac…
Re: Experiencing decreased performance with ChatGPT-4
#95Earlier quoted context omitted.
We will never know for sure, it is equally likely they did some cost savings which caused a reduction in quality. I certainly noticed that too, but we have no way to prove that, any proof can always be dismissed easily. For example I see it generate code with hallucinated variables quite often now, that never happened before to that degree, but I'm just as well part of the group hallucinating so easy to dismiss. Anec…
> We will never know for sure, it is equally likely they did some cost savings which caused a reduction in quality. That is entirely not equally likely, and would be completely unprecedented, at the frontier of an emerging technology that people are pumping the money and the future of the world into to win.
Combined with stricter guardrails, I would certainly expect intelligence to go down.
Check out the tikz unicorn drawing example from the paper "Sparks of Artificial General Intelligence: Early experiments with GPT-4" (pdf page 7)
Re: Experiencing decreased performance with ChatGPT-4
#96Without concrete examples, I do wonder if much of this is perceptual. I love using ChatGPT, but once the amazement that it works as well as it does has worn off, one ends up spotting the flaws more than before. I feel that advocates and critics of ChatGPT are both right, to a degree, but looking at the models responses from slightly different angles: it wouldn't be surprising if users' angles shift over time.
Re: Experiencing decreased performance with ChatGPT-4
#97I don't know about the API, but the ChatGPT UI GPT4 really became much worse, so much so that I had to cancel my subscription. It's not just about novelty factor. I used to store all my prompts locally, and when I compare old responses and new ones, there is a huge difference. OpenAI employee said the API model doesn't change, but were careful to not say anything about the UI. I am now waiting to get access to the AP…
At this point the default assumption should be "people who are downvoting posts like these are dis-ingenious and do so with an agenda". Otherwise, I had to do the same, as it has been so lobotomized that doing many tasks I used to delegate to GPT-4 has become easier to do manually once again.
Re: Experiencing decreased performance with ChatGPT-4
#98I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.
The same exact prompts from February are leading to significantly degraded responses. I've proven it to myself dozens of times.
Re: Experiencing decreased performance with ChatGPT-4
#99Earlier quoted context omitted.
We will never know for sure, it is equally likely they did some cost savings which caused a reduction in quality. I certainly noticed that too, but we have no way to prove that, any proof can always be dismissed easily. For example I see it generate code with hallucinated variables quite often now, that never happened before to that degree, but I'm just as well part of the group hallucinating so easy to dismiss. Anec…
> We will never know for sure, it is equally likely they did some cost savings which caused a reduction in quality. That is entirely not equally likely, and would be completely unprecedented, at the frontier of an emerging technology that people are pumping the money and the future of the world into to win.