I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.
Experiencing decreased performance with ChatGPT-4
41–50 of 200 posts
Re: Experiencing decreased performance with ChatGPT-4
#42Earlier quoted context omitted.
They've definitely changed something about the models, and it is in their interests to do so, both to create a low-latency experience, but most importantly, to save money. While GPT-4 is still workable, GPT-3.5 flatly refuses requests these days, claiming that as an "AI language model" it couldn't help me write code.
Did you get the help you needed in the end? I've seen it do that once but was able to cajole it within 1-2 prompts.
Re: Experiencing decreased performance with ChatGPT-4
#43While calling an API is convenient (and with the implicit promise of always performing to a certain standard) this is anything but guaranteed
The examples on the thread are interesting, I wonder if the wording might have changed slightly or if the 'human fine tuning" loop might introduce certain instabilities in some specific tasks
Re: Experiencing decreased performance with ChatGPT-4
#44I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.
AFAIK, OpenAI has repeatedly stated that GPT4 hasn't changed. People repeatedly states that when they use ChatGPT, they get a difference experience today than before. Both can be true at the same time, as ChatGPT is a "packaged" experience of GPT4, so if you use the API versions, nothing has likely changed. But ChatGPT has guaranteed changed, for better or worse, as that's "just" integration work rather than fundamental changes to the model.
In the discussions on HN, people tend to speak each other regarding this as well, saying things like "GPT4 has for sure changed" when their only experience of GPT4 is via ChatGPT, which has changed since launch, obviously.
But ChatGPT != GPT4, which could always be made clearer.
Re: Experiencing decreased performance with ChatGPT-4
#45---
Failure
Q: How can you use a 4-minute hourglass and a 7-minute hourglass to measure 9 minutes?
A: Here's a way to measure exactly 9 minutes using a 4-minute hourglass and a 7-minute hourglass:
Start both hourglasses at the same time. When the 4-minute hourglass runs out, flip it over (4 minutes elapsed). When the 7-minute hourglass runs out, flip the 4-minute hourglass again (7 minutes elapsed). The 4-minute hourglass had 1 minute left when you flipped it, so it will run for an additional minute. When the 4-minute hourglass runs out this time, 2 additional minutes have passed (for a total of 9 minutes elapsed). So, you've just measured exactly 9 minutes using a 4-minute hourglass and a 7-minute hourglass.
---
Success
Q: How can you use a 4-minute hourglass and a 7-minute hourglass to measure 9 minutes? Please reason step-by-step. You are allowed to run preparatory steps beforehand if necessary.
A: The problem can be solved by using the following steps:
Preparatory Steps:
Start both hourglasses at the same time. When the 4-minute hourglass runs out, turn it over to start it again. (This marks 4 minutes.) When the 7-minute hourglass runs out, don't flip it yet, but note that the 4-minute hourglass has now been running for 3 minutes on its second run. (This marks 7 minutes.) When the 4-minute hourglass runs out again, flip the 7-minute hourglass. (This marks 8 minutes.)
Measurement Steps:
When the sand in the 7-minute hourglass runs out again, it will have been 1 minute since its flip, marking a total of 9 minutes. So, the process will be: 0-4 minutes: first run of the 4-minute hourglass 4-7 minutes: second run of the 4-minute hourglass until the 7-minute hourglass runs out 7-8 minutes: rest of the second run of the 4-minute hourglass 8-9 minutes: one minute of the 7-minute hourglass after flipping it when the 4-minute hourglass runs out the second time.
Thus, the total time measured is 9 minutes.
Re: Experiencing decreased performance with ChatGPT-4
#46I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.
We also can never independently evaluate it, OpenAI could cache messages, fine tune on public test sets, etc. etc.
Re: Experiencing decreased performance with ChatGPT-4
#47I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.
I remember the first time I played Minecraft and I was in awe at how expansive the play world felt. Without thinking too much about it, I had the feeling that if I set off in any direction I would discover infinitely new things. After enough playtime I saw the repeating patterns and eventually it felt so small again.
Re: Experiencing decreased performance with ChatGPT-4
#48I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.
Re: Experiencing decreased performance with ChatGPT-4
#49Re: Experiencing decreased performance with ChatGPT-4
#50I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.
We will never know for sure, it is equally likely they did some cost savings which caused a reduction in quality. I certainly noticed that too, but we have no way to prove that, any proof can always be dismissed easily. For example I see it generate code with hallucinated variables quite often now, that never happened before to that degree, but I'm just as well part of the group hallucinating so easy to dismiss. Anec…
That is entirely not equally likely, and would be completely unprecedented, at the frontier of an emerging technology that people are pumping the money and the future of the world into to win.