Live data from Hacker News

Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?

news.ycombinator.com

1–10 of 44 posts

Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?

#1
Has there been any consensus on this phenomenon I've seen where there are reports of decrease in model performance? This decrease in quality seems to take many forms: laziness, lower expressiveness, mistakes, laziness etc.

One the other hand there are people who claim there hasn't been any degradation at all[0]?

If there is indeed no degradation how could the perceived degradation be explained?

[0] https://community.openai.com/t/declining-quality-of-openai-m...

https://community.openai.com/t/declining-quality-of-openai-m...

https://www.reddit.com/r/OpenAI/comments/18sc92o/with_all_th...

http://arxiv.org/pdf/2307.09009

Re: Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?

#4
> If there is indeed no degradation how could the perceived degradation be explained?

By being disproportionately impressed previously. Maybe in the early days people were so impressed by their little play experiments they forgave the shortcomings. Now that the novelty is wearing off and they try to use it for productive work, the scales tipped and failures are given more weight.

Re: Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?

#5
post #3

Does it matter? Perhaps a more important question is whether there has been a decline in user satisfaction.

A user can be satisfied because they're receiving an answer, but it might be a totally wrong answer, maybe in some obvious way like 2+2=5, or in a much less obvious way, like generating a /mostly/ accurate biography of a person which includes a year long period of their life that never happened. There needs to be a measurable criteria to judge performance over a variety of metrics over time, because what we notice on a surface level while interacting with AI might not represent the actual performance of accuracy of the outputs we're receiving.

Re: Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?

#6
One of the things I've personally observed is that ChatGPT has become very verbose these days. Previously, it used to return the right amount of information in most contexts, and I can't get that behavior back with prompts asking it to be concise, because then it'll just omit important parts, prioritizing providing a extremely high-level summary that elucidates very little.

No opinion on Claude because I've not had a long experience using it, but as it stands Claude 3 Sonnet is usually better at inferring what's asked of it over ChatGPT.

Re: Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?

#7

One of the things I've personally observed is that ChatGPT has become very verbose these days. Previously, it used to return the right amount of information in most contexts, and I can't get that behavior back with prompts asking it to be concise, because then it'll just omit important parts, prioritizing providing a extremely high-level summary that elucidates very little. No opinion on Claude because I've not had a…

+1 on verbosity - it happened when switching from 4t to 4o I think, and personally I don’t like it.

Should be fizable with system prompt though.

Re: Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?

#9

One of the things I've personally observed is that ChatGPT has become very verbose these days. Previously, it used to return the right amount of information in most contexts, and I can't get that behavior back with prompts asking it to be concise, because then it'll just omit important parts, prioritizing providing a extremely high-level summary that elucidates very little. No opinion on Claude because I've not had a…

Agreed on both counts. I have stopped using ChatGPT due to its verbosity, bloating the price. Despite the prompt I cannot get it to cut to the chase. Claude has been much better in this regard. There was a very distinct change in ChatGPT behavior even using the same model towards verbosity. The cynic in me supposes it’s to bloat revenue.

Re: Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?

#10
post #5
post #3

Does it matter? Perhaps a more important question is whether there has been a decline in user satisfaction.

A user can be satisfied because they're receiving an answer, but it might be a totally wrong answer, maybe in some obvious way like 2+2=5, or in a much less obvious way, like generating a /mostly/ accurate biography of a person which includes a year long period of their life that never happened. There needs to be a measurable criteria to judge performance over a variety of metrics over time, because what we notice on…

Ok, "does it matter" was hyperbole.

But if users are getting less satisfied that's bad news for OpenAI, whether the quality is objectively worse or not.

Post reply on HN