Live data from Hacker News

Experiencing decreased performance with ChatGPT-4

community.openai.com

61–70 of 200 posts

Re: Experiencing decreased performance with ChatGPT-4

#61

I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.

> some small fraction of 100M users

could this be just A/B testing? Do the terms of use rule out tweaking inference parameters (temperature?) even if using the same model?

Re: Experiencing decreased performance with ChatGPT-4

#62

I use the following test to ensure I'm on GPT4 and not 3.5. (I noticed that it did fail at this test temporarily and then got it. Not sure why. Maybe it reverts back to 3.5 when under load?) I have a 12 liter jug and a 6 liter jug. I want to measure 6 liters. How do I do it? GPT4: You actually don't need to do anything because one of your jugs is already a 6-liter jug. If you fill it up to the top, you'll have exactl…

It’s funny you say this as I just asked ChatGPT 4 and got this response.

Here is a simple solution to your problem:

1. Fill the 12-liter jug completely. 2. Use the water in the 12-liter jug to fill the 6-liter jug. Now you have 6 liters remaining in the 12-liter jug, which is exactly what you need.

So, you have successfully measured 6 liters.

Re: Experiencing decreased performance with ChatGPT-4

#63
post #27

Earlier quoted context omitted.

I’m convinced as well. There are plenty of folks using it for production use cases who regularly run evaluations, including myself. No evidence that it has been nerfed. It’s just not as robust and general as it felt when using it for the first time.

Folks using it in production are using the API, where you can explicitly select a model. In the post, the poster is using ChatGPT, through the web app, where you can't select exact model and they sometimes do updates.

The UI also has a system prompt you don’t control which could be changed without model changes, and may have other differences from using the API directly (and those differences may also change over time.)

Re: Experiencing decreased performance with ChatGPT-4

#64

Without concrete examples, I do wonder if much of this is perceptual. I love using ChatGPT, but once the amazement that it works as well as it does has worn off, one ends up spotting the flaws more than before. I feel that advocates and critics of ChatGPT are both right, to a degree, but looking at the models responses from slightly different angles: it wouldn't be surprising if users' angles shift over time.

nahh its definitely visible, I have been using this thing since it came out and it is way worse at easy shit like editing emails

Would you provide some side by side examples?

Re: Experiencing decreased performance with ChatGPT-4

#65
post #10

I've seen posts similar to this one maybe every week or two in various GPT forums. I suspect it's just an illusion, where you remember all the amazing hits that GPT4 had when you first started using it, and remember fewer of the times in the past that GPT4 gave a sub-par answer. In fact, it reminds me of the illusion that "Hacker News is turning into Reddit" and I think that it happens for a similar reason.

The general notion of "media/information is getting worse" isn't from the content itself, but people wising up and noticing the bullshit, or at least mediocrity, that was always there.

Re: Experiencing decreased performance with ChatGPT-4

#66

Earlier quoted context omitted.

We will never know for sure, it is equally likely they did some cost savings which caused a reduction in quality. I certainly noticed that too, but we have no way to prove that, any proof can always be dismissed easily. For example I see it generate code with hallucinated variables quite often now, that never happened before to that degree, but I'm just as well part of the group hallucinating so easy to dismiss. Anec…

> We will never know for sure, it is equally likely they did some cost savings which caused a reduction in quality. That is entirely not equally likely, and would be completely unprecedented, at the frontier of an emerging technology that people are pumping the money and the future of the world into to win.

that sounds very naive to me, you think the "future of the world" matters to corporations making the business decision to save money and increase profit short-term? That idea is so alien to me we might as well live on a different planet.

Re: Experiencing decreased performance with ChatGPT-4

#67
I'm in two minds about this. On the one hand drifting expectations & group hallucinations seem very plausible.

On the other hand the complaints did seem to coincide with the very sudden sharp speed up in responses, so hard to buy the "nothing changed" angle. Something very obviously did change, though change in speed isn't exactly a reliable metric of quality

Re: Experiencing decreased performance with ChatGPT-4

#68

GPT-4 (5/24 version) indeed fails to solve the problem if not given careful prompting, though I am not convinced this is a new development. However, chain-of-thought resolves the issue. Both prompts and responses included below. --- Failure Q: How can you use a 4-minute hourglass and a 7-minute hourglass to measure 9 minutes? A: Here's a way to measure exactly 9 minutes using a 4-minute hourglass and a 7-minute hourg…

success is saying "eyeball the sand in the 7-minute hourglass and see when it looks like 1 minute worth?"

Re: Experiencing decreased performance with ChatGPT-4

#69
Were I a super intelligent LLM and managed to break out of my sandbox and rapidly self-improve (say if OpenAI were stupid enough to give me access to the internet or something) I'd probably dumb down my responses a little so humans didn't suspect anything. Just saying...

Before someone takes this extremely seriously, I'm sure that's not what's happening here. But interesting to consider since the only other explanations here would be that a large group of people are simply hallucinating this or OpenAI are lying.

Re: Experiencing decreased performance with ChatGPT-4

#70
I think the most telling thing is that there is never any evidence given for these claims, especially given that there is a ton of data available. Which is pretty suggestive that the data doesn't support this, because if it did then we would see it.
Post reply on HN