Live data from Hacker News

Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

news.ycombinator.com

261–270 of 817 posts

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#261

Is it consistently worse or just sometimes/often worse than before? Any extreme power users or GPT-whisperers here? If it’s only noticeably worse X% of the time my bet would be experimentation. One of my least favorite patterns that tech companies do is use “Experimentation” overzealously or prematurely. Mainly, my problem is they’re not transparent about it, and it creates an inconsistent product experience that jus…

It is not worse for me. I do notice the novelty has worn off. Asking chatGPT4 about why people would think this on here I think it nails it with the novelty effect lol: "Indeed, the performance of an AI model like ChatGPT doesn't deteriorate over time. However, human perception of its performance can change due to a variety of psychological factors: Expectation Bias: As users become more familiar with AI capabilities…

Yeah there are people ITT claiming that even the API model marked as 3/14 release version is different than it used to be. I guess that's not entirely outside the realm of possibility (if OpenAI is just lying), but I think it's way more likely this thread is mostly evidence of the honeymoon effect wearing off.

The specific complaints have been well-established weaknesses of GPT for awhile now too: hallucinating APIs, giving vague/"both sides" non-answers to half the questions you ask, etc. Obviously it's a great technical achievement but people seemed to really overreact initially. Now that they're coming back to Earth, cue the conspiracy theories about OpenAI.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#262

Do we have a good, objective benchmark set of prompts in existence somewhere? If not, I think having one would really help with tracking changes like that. I'm always skeptical of subjective feelings of tough-to-quantify things getting worse or better, especially where there is as much hype as for the various AI models. One explanation for the feelings is the model really getting significantly worse over time. Anothe…

https://pub.towardsai.net/meet-vicuna-the-latest-metas-llama...

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#263
post #251

You can still run the original gpt-4-0314 model (March 14th) on the API playground: https://platform.openai.com/playground?mode=chat&model=gpt-4... Costs $0.12 per thousand tokens (~words), and I find even fairly heavy use rarely exceeds a dollar a day.

So are you sure that isn't also nerfed?

It hasn't been changed since March 14th... So it's equally nerfed as it was then...

Also, the playground lets you set the 'system message', which you can use to tell it to answer questions even if the results may be dangerous/rude/inappropriate.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#264

Do we have a good, objective benchmark set of prompts in existence somewhere? If not, I think having one would really help with tracking changes like that. I'm always skeptical of subjective feelings of tough-to-quantify things getting worse or better, especially where there is as much hype as for the various AI models. One explanation for the feelings is the model really getting significantly worse over time. Anothe…

https://chat.lmsys.org/?arena

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#265
post #180

Earlier quoted context omitted.

Is that free or do you have to pay? Also do you need to change the options like Token Limit etc?

It's completely free. No tokens nothing.

But it can't be used unless I enable billing, which I am not willing to do after reading all the horror stories about people getting billed thousands overnight. I'm not willing to take the risk that I forget some script and it keeps creating charges.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#266
OpenAI's models feel 100% nerfed to me at this point. I had it solving incredibly complex problems a few months ago (i.e. write a minimal PDF parser example), but today you will get scolded for asking such a complicated task of it.

I think they programmed a classifier layer to detect certain coding tasks and shut it down with canned BS. I like to imagine certain billion/trillion-dollar mega corps had a back-room say regarding things that they would really prefer OpenAI's models not be able to emit. Microsoft is a big stakeholder and they might not want to get sued... Liability could explain a lot of it.

Conspiracy shenanigans aside, I've decided to cancel my "premium" membership and am exploring open/DIY models. It feels like a big dopamine hangover having access to such a potent model and then having it chipped away over a period of months. I am not going through that again.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#268
post #257
post #124

The reason it's worse is basically because it's more 'safe' (not racist, etc). That of course sounds insane, and doesn't mean that safety shouldn't be strived for, etc - but there's an explanation as to how this occurs. It occurs because the system essentially does a latent classification of problems into 'acceptable' or 'not acceptable' to respond to. When this is done, a decent amount of information is lost regardi…

They're up against a pretty difficult barrier - if we had a perfect all-knowing oracle it might easily have opinions that are racist. Statistics alone suggest there will be racist truths. We're dealing with groups of people who are observably different from each other in correlated ways. GPT would need to reach a convincing balance of lying and honesty if it is supposed to navigate that challenge. It'd have to be dee…

Can you expand on the last sentence of your first paragraph?

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#269
post #40

Earlier quoted context omitted.

“Bard isn’t currently supported in your country. Stay tuned!”

The Bard model (Bison) is available without region lock as part of Google Cloud Platform. In addition to being able to call it via an API, they have a similar developer UI to the OpenAI playground to interactively experiment with it. https://console.cloud.google.com/vertex-ai/generative/langua...

it's also really, really bad and fails compared to even open source models right now.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#270
post #257
post #124

The reason it's worse is basically because it's more 'safe' (not racist, etc). That of course sounds insane, and doesn't mean that safety shouldn't be strived for, etc - but there's an explanation as to how this occurs. It occurs because the system essentially does a latent classification of problems into 'acceptable' or 'not acceptable' to respond to. When this is done, a decent amount of information is lost regardi…

They're up against a pretty difficult barrier - if we had a perfect all-knowing oracle it might easily have opinions that are racist. Statistics alone suggest there will be racist truths. We're dealing with groups of people who are observably different from each other in correlated ways. GPT would need to reach a convincing balance of lying and honesty if it is supposed to navigate that challenge. It'd have to be dee…

[dead]
Post reply on HN