Live data from Hacker News

Experiencing decreased performance with ChatGPT-4

community.openai.com

101–110 of 200 posts

Re: Experiencing decreased performance with ChatGPT-4

#101

I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.

The quality seriously degraded overnight a few months ago... it was quite abrupt and obvious to those who use it regularly.

It's not surprising, they were probably running the model at a huge and ultimately unacceptable loss. But they should really offer a higher paid tier to access the previous capabilities... not drop them entirely. Many would pay far more than $20/month to access a marginally but meaningfully better model.

EDIT: Many being dismissive of LLMs don't even seem to use them. Providers are vastly overvalued from an investment perspective, but the utility is very real. To say the loss in capability is just an "illusion" is clearly wrong to anybody who actually uses it.

Re: Experiencing decreased performance with ChatGPT-4

#102

I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.

We will never know for sure, it is equally likely they did some cost savings which caused a reduction in quality. I certainly noticed that too, but we have no way to prove that, any proof can always be dismissed easily. For example I see it generate code with hallucinated variables quite often now, that never happened before to that degree, but I'm just as well part of the group hallucinating so easy to dismiss. Anec…

This is a common tactic observed with toxic personality disorders. People will repeatedly ask for examples knowing they can dispute any example given because the topic is subjective. You can spot the loop in these threads. An army of people comment/flag asking for examples regardless of how many are given in the thread. When you provide them examples they nitpick and call your prompting bad. Not saying it's bots but there's a pattern with these "OpenAI nerfed" threads across social media right now.

Re: Experiencing decreased performance with ChatGPT-4

#103

Earlier quoted context omitted.

We will never know for sure, it is equally likely they did some cost savings which caused a reduction in quality. I certainly noticed that too, but we have no way to prove that, any proof can always be dismissed easily. For example I see it generate code with hallucinated variables quite often now, that never happened before to that degree, but I'm just as well part of the group hallucinating so easy to dismiss. Anec…

> We will never know for sure, it is equally likely they did some cost savings which caused a reduction in quality. That is entirely not equally likely, and would be completely unprecedented, at the frontier of an emerging technology that people are pumping the money and the future of the world into to win.

A technology that is expensive to run and hard to scale. They’re doing work trying to scale, in what world is this unprecedented?

Re: Experiencing decreased performance with ChatGPT-4

#104

I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.

I've run the same prompts as before and received different responses. I'm not sure how that's a hallucination.

Re: Experiencing decreased performance with ChatGPT-4

#105

I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.

Slowly but surely, the comment gaslighting all of the people reporting the issue, makes its way to the top, while other comments with genuine discussion are flagged and slip lower. Seen this before...

Re: Experiencing decreased performance with ChatGPT-4

#106
post #82

I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.

My suspicion is that we're collectively becoming accustomed to ChatGPT failures. These failures cause problems, and become more annoying with time. The same thing happened with voice assistants. That being said, the safety filters have definitively changed in OpenAI. ChatGPT is definitely more prone to reminding me that it is an LLM, and it refuses to participate in pretend play which it perceives as violating its sa…

The filters really have changed.

I started using it relatively late, but earlier in May, you could have given it a DOI link, and it would have summarized it for you. Now, it argues that it's not a database and that it can only summarize it if you provide the full text. However, if you ask for it with the title of the paper, it will provide you with a summary.

You could have also asked it to search patents on some topic, and it would have given you a list of links. Now, it provides instructions on how to find it yourself.

Re: Experiencing decreased performance with ChatGPT-4

#107

Earlier quoted context omitted.

I think it's more likely that people are confused, and OpenAI is not making things any clearer either. AFAIK, OpenAI has repeatedly stated that GPT4 hasn't changed. People repeatedly states that when they use ChatGPT, they get a difference experience today than before. Both can be true at the same time, as ChatGPT is a "packaged" experience of GPT4, so if you use the API versions, nothing has likely changed. But Chat…

I’d note they explicitly document they rev GPT-4 every two weeks and provide fixed snapshots of the prior periods model for reference. One could reasonably benchmark the evolution of the model performance and publish the results. But certainly you’re right - ChatGPT != GPT4, and I would expect that ChatGPT performs worse than GPT4 as it’s likely extremely constrained in its guidance, tunings, and whatever else they d…

Heh, this kind of reminds me of the process of enterprise support.

Working with the customer in dev: "Ok, run this SQL query and restart the service. Done, ok does the test case pass?" Done in 15 minutes.

Working with customer in production: "Ok, here is a 35 point checklist of what's needed to run the SQL query and restart the service. Have your compliance officer check it and get VP approval, then we'll run implementation testing and verification" --same query and restart now takes 6 hours.

Re: Experiencing decreased performance with ChatGPT-4

#108

Notice how you never hear anyone saying that GPT-4 is better since the launch. You'd expect to hear something like that as people gain more experience with prompting it. I've certainly noticed that the quality of responses has gone down, and I have to repeat myself more often as it doesn't always remember all my instructions. For an example of something it can no longer do, I used to show it off by having it explain…

> Notice how you never hear anyone saying that GPT-4 is better since the launch. You'd expect to hear something like that as people gain more experience with prompting it.

I'd expect the opposite. The first time you use ChatGPT (or GPT-4), you're in awe of what it can do, and more willing to overlook failures. As you use it, it becomes more mundane, and the instances where it messes up become more obvious.

Re: Experiencing decreased performance with ChatGPT-4

#109

I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.

I use the API, not the chat site. Since 30 June, the API responses are making common English misspelling errors, of the type where two words sound the same with different meanings such as break and brake. I saw this happen zero times in the prior GPT-4 model, and multiple times this July, on multiple conversation topics and multiple word pairs. Curiously, they're behaving as misspellings rather than mismeanings, sinc…

Or https://en.wikipedia.org/wiki/Flowers_for_Algernon

Re: Experiencing decreased performance with ChatGPT-4

#110

GPT-4 (5/24 version) indeed fails to solve the problem if not given careful prompting, though I am not convinced this is a new development. However, chain-of-thought resolves the issue. Both prompts and responses included below. --- Failure Q: How can you use a 4-minute hourglass and a 7-minute hourglass to measure 9 minutes? A: Here's a way to measure exactly 9 minutes using a 4-minute hourglass and a 7-minute hourg…

success is saying "eyeball the sand in the 7-minute hourglass and see when it looks like 1 minute worth?"

It's just missing a few steps: First, remove the end from the lower (still empty) half of 7-minute hourglass, causing it to drain its sand, while it runs. When the 7-minute hourglass runs out, do the same with the 4-minute hourglass (being on its second run). Flip the (now empty) 7-minute hourglass over and let the rest of the 4-minute hourglass drain into this. As the latter runs out, flip the 7-minute hourglass for 1 minute worth of accumulated sand (minus a few grains lost over handling procedures). ;-)
Post reply on HN