Live data from Hacker News

Experiencing decreased performance with ChatGPT-4

community.openai.com

121–130 of 200 posts

Re: Experiencing decreased performance with ChatGPT-4

#121

I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.

I am pretty sure it is not. I have posted a concrete example of GPT-Copilot enshittification here: https://github.com/orgs/community/discussions/60116

Re: Experiencing decreased performance with ChatGPT-4

#122

I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.

I think it's more likely that people are confused, and OpenAI is not making things any clearer either. AFAIK, OpenAI has repeatedly stated that GPT4 hasn't changed. People repeatedly states that when they use ChatGPT, they get a difference experience today than before. Both can be true at the same time, as ChatGPT is a "packaged" experience of GPT4, so if you use the API versions, nothing has likely changed. But Chat…

>But ChatGPT != GPT4, which could always be made clearer.

Isn't the thread about ChatGPT? I mean it is helpful to know that they are not the same (I personally was not clear on this myself, so I, at least, benefitted from your comment), but I think the thread is just about Chat GPT.

Re: Experiencing decreased performance with ChatGPT-4

#123

Anecdotal: I introduced my doctor to ChatGPT and Bard many months ago and they were impressed. Fast forward a few days ago and I asked them if they had used either since. They said it was far inferior to Google, so no. So I asked them to show me an example. Basically any medical question was answered with “go ask a doctor”. I suppose because of liability concerns. Both were basically useless. So this decreased perfor…

I don't understand your anecdote. I'm able to ask it medical questions and get answers, for example: https://chat.openai.com/share/75f94000-552f-42d6-aadf-198fd9... https://chat.openai.com/share/0933abf7-1015-41b5-9a49-ca2b6e... Whether someone should trust the answers is a different question.

He asked a dosage question similar to your second example and without tricking it, it would not give a response. The drugs were not as common (at least I wasn’t familiar with them) as what you listed, so maybe that had something to do with it.

Re: Experiencing decreased performance with ChatGPT-4

#124

I’m convinced this is group hallucination. It must be so interesting to work at OpenAI, knowing you didn’t change a thing, and seeing that because of random chance, some small fraction of 100M users have all tricked each other that suddenly, something is different.

> I’m convinced this is group hallucination

Can you provide some evidence to back that up? Especially because OpenAI _has_ been tinkering with ChatGPT - by trying to limit jailbreaks.

People have a strong prior that these kinds of changes will reduce model performance (because you're limiting your model), so the burden is on you to show that performance hasn't degraded.

Re: Experiencing decreased performance with ChatGPT-4

#125
After using the API quite heavily, I believe both arguments (expectations changed vs. ChatGPT changed) are likely correct.

As with any technology, we should predict the novelty to wear off and the rough edges to become more apparent. Peoples’ expectations have changed.

On the flip side, the ChatGPT interface has also changed (regardless of the underlying model). Any context you add to an LLM prompt will steer the LLM’s output, better or not.

We know for a fact ChatGPT uses a different/additional prompt to the API, as ChatGPT always has the current date. This changed in the ChatGPT interface around May 13th, around the same time as OpenAI claimed the model was identical. The addition of Plugins/web browsing around that time also made it easier to pollute your prompt, if either were enabled.

We also know that ChatGPT is running on different infra (based on latency diffs), so even if it’s the same model, it’s possible it’s configured ever so differently.

And finally, we also know there’s a new model (woohoo function calling!).

As an API user, my personal experience with GPT-4 is very similar to when it first came out. The hype was very high (AGI in months!) and the reality has been quite different.

GPT-4 is an amazing mirror of society, but it’s only worth what you put in.

Re: Experiencing decreased performance with ChatGPT-4

#126
post #115

Earlier quoted context omitted.

Would you provide some side by side examples?

This is the strongest point of evidence I have that the phenomenon isn't real - One can very easily recreate prompts and share two links from different eras, yet we never see that. My guess is that the complainers spent a lot of time finding narrow queries that worked once and now, the horrors of stochasticity are breaking their ability to recreate those narrow queries for new topics. Kind of a different flavor to al…

I’m thinking the same thing. I have the ChatGPT app and play around with it. It saves all queries and responses. It would be trivial to copy and paste and recreate to show proof.

In fact, I took a query from a few months ago which was a trick question and reran it and got effectively the same, correct answer.

Re: Experiencing decreased performance with ChatGPT-4

#127
We’ve been testing the upgraded models in the API (where you can control when the upgrade happens), and the newer ones perform significantly worse than the older ones on the same tasks. Tweaking the prompts helps some but not enough. We’re staying on the older models for now in production.

Hope OpenAI figures this out because quality has been their biggest moat up until now.

Re: Experiencing decreased performance with ChatGPT-4

#128
I’m also of the opinion my paid chatgpt responses have been lower quality.

However users going on multi page rants is so bad. I read 30% through and its like “dude stop”

Its a product and you pay for it, if its not working then express the feedback and move on. Making demands of OpenAI like it is some elected government agency with transparency requirements is outrageous

Re: Experiencing decreased performance with ChatGPT-4

#129
post #75

Earlier quoted context omitted.

I think it's more likely that people are confused, and OpenAI is not making things any clearer either. AFAIK, OpenAI has repeatedly stated that GPT4 hasn't changed. People repeatedly states that when they use ChatGPT, they get a difference experience today than before. Both can be true at the same time, as ChatGPT is a "packaged" experience of GPT4, so if you use the API versions, nothing has likely changed. But Chat…

It's a bit of both. The GPT-4 models have definitely been changing - there's multiple versions right now and you can try them out in the Playground. One of the biggest differences is that the latest model patches all of the GPT-4 jailbreak prompts; quite a big change if you were doing anything remotely spicy. But OA also says that it hasn't been changing the underlying model beyond that (that's probably the tweet you…

It'd be insane if OpenAI wasn't changing GPT-4. That kind of flat footedness would cost them their entire first mover advantage.

Re: Experiencing decreased performance with ChatGPT-4

#130

GPT-4 (5/24 version) indeed fails to solve the problem if not given careful prompting, though I am not convinced this is a new development. However, chain-of-thought resolves the issue. Both prompts and responses included below. --- Failure Q: How can you use a 4-minute hourglass and a 7-minute hourglass to measure 9 minutes? A: Here's a way to measure exactly 9 minutes using a 4-minute hourglass and a 7-minute hourg…

success is saying "eyeball the sand in the 7-minute hourglass and see when it looks like 1 minute worth?"

No, this is precisely correct. To simplify further:

1. Start H4 and H7

2. Flip H4 when it runs out (4-minute mark)

3. Flip H7 when it runs out (7-minute mark, 1 minute left on H4)

4. Flip H7 back when H4 runs out again (8-minute mark, 1 minute elapsed on H7)

5. When H7 runs out again (after 1 minute), exactly 9 minutes have passed

Post reply on HN