Live data from Hacker News

Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

news.ycombinator.com

121–130 of 817 posts

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#121

Earlier quoted context omitted.

FWIW, I started to get the same feeling as the OP about GPT-4 model I have access to on Azure, so if there's any deal being cut here, it might involve dumbing down the model for paying Azure customers as well. Now, to be clear: I only started to get a feeling that GPT-4 on Azure is getting worse. I didn't do any specific testing for this so far, as I thought I may just be imagining it. This thread is starting to conv…

I’ve seen degradation in the app and via the API, so if I had to bet, they’ve probably kneecapped the model so that it works passably everywhere they’ve been made it available vs. works well in one place or another.

Yes. I think 'sirsinsalot is likely right in suggesting[0] that they could be trying "to hoard the capability to out compete any competitor, of any kind, commercially or politically and hide the true extent of your capability to avoid scrutiny and legislation", and that they're currently "dialing back the public expectations", possibly while "deploying the capability in a novel way to exploit it as the largest lever" they can.

That view is consistent with GPT-4 getting dumber on both OpenAI proper and Azure OpenAI - even as the companies and corporations using the latter are paying through the nose for the privilege.

Alternative take is that they're doing it to slow the development of the whole field down, per all the AI safety letters and manifestos that they've been signing and circulating - but that would be at best a stop-gap before OSS models catch up, and it's more than likely that OpenAI and/or Microsoft would succumb to the temptation of doing what 'sirsinsalot suggested anyway.

--

[0] - https://news.ycombinator.com/item?id=36135425

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#122

It's time to design a public benchmark for these types of systems to compare between versions. Of course, any vendor who trains on the benchmark should face extreme contempt, but we'd also need to generate novel questions of equal complexity. Alternatively, there should be a trusted auditor who uses a secret benchmark.

But this is the same version that changes without a change of the version number.

Well, people suspect it isn't, and it's not like we can see the internal version designation, and it's not even like we would care a lot, if it performed identically from day to day.

Indeed, you could do better or worse with the exact same raw checkpoint, just depending on inference-optimizing tricks.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#123
post #89

Earlier quoted context omitted.

Not necessarily American, you just have to avoid EU and, I believe, Russia/China/Cuba etc.

I'm in Switzerland and Bard is locked out, we do not go by EU laws because we are not part of the EU. We have plenty of bilateral deals but still.

In practice Switzerland adopts EU law with minor revisions because doing otherwise would lock Swiss businesses out of the EU internal market.

The Swiss version of GDPR is coming in September:

https://www.ey.com/en_ch/law/a-new-era-for-data-protection-i...

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#124
The reason it's worse is basically because it's more 'safe' (not racist, etc). That of course sounds insane, and doesn't mean that safety shouldn't be strived for, etc - but there's an explanation as to how this occurs.

It occurs because the system essentially does a latent classification of problems into 'acceptable' or 'not acceptable' to respond to. When this is done, a decent amount of information is lost regarding how to represent these latent spaces that may be completely unrelated (making nefarious materials, or spouting hate speech are now in the same 'bucket' for the decoder).

This degradation was observed quite early on with the tikz unicorn benchmark, which improved with training, and then degraded when fine-tuning to be more safe was applied.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#125

Maybe you annoyed it? I’m super nice to it and it performa better than ever. I ask it lay-of-the-land questions about technical problems that are new to me to detailed coding problems that I understand well but don’t want to figure out. The best though, is helping me navigate complicated UI’s. I tell it how I want some complicated software / website to behave, and it’ll tell me the arcane menu path to follow. It’s fu…

This might seem funny, but I noticed this too.

When I thank GPT-4 and also give it feedback on what worked and what didn't, it works better and so to me it seems like the equivalent of "putting in more effort".

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#126
post #84

Given the incoming compute capability from nvidia and the speed of advancement, we have to stop and think ... does it make sense to give access, paid or otherwise, to these models once they reach a certain sophistication? Or does it make even more sense to hoard the capability to out compete any competitor, of any kind, commercially or politically and hide the true extent of your capability to avoid scrutiny and legi…

I have some first hand thoughts. I think overall the quality is significantly poorer on GPT4 with plugins and bing browsing enabled. If you disable those, I am able to get the same quality as before. The outputs are dramatically different. Would love to hear what everyone else sees when they try the same.

would be alarming if you had second hand thoughts...

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#127

There's no doubt that it's gotten a lot worse on coding, I've been using this benchmark on each new version of GPT-4 "Write a tiptap extension that toggles classes" and so far it's gotten it right every time, but not any more, now it hallucinates a simplified solution that don't even use the tiptap api any more. It's also 200% more verbose in explaining it's reasoning, even if that reasoning makes no sense whatsoever…

It was a great ride while it lasted. My assumption is that efficacy at coding tasks is such a small percent of users, they’ve just sacrificed it on the altar of efficiency and/or scale. That, or they’ve cut some back room deal with Microsoft to make Copilot have access to the only version of the model that can actually code.

Maybe it had to do with jailbreaks? A lot of the jailbreaks were related to coding, so maybe they put more restrictions in there. Only speculating, but I cannot imagine why it got worse otherwise.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#128

This is the normal workflow for drug dealers too. The first fix is free. The second one will cost you money. The third one will be laced with fillers and have degraded quality.

But I don't get why they would degrade quality? Maybe compression to save resources? Otherwise I wouldn't know what they have to gain from it. If anything, they might lose paid subscribers. I loved the results for the $20 I spent, now I'm not so sure anymore.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#129

Earlier quoted context omitted.

Yes, that was how I read it as well. I was just pointing out that the public release was already extremely nerfed from what was available pre-launch.

Interesting, please expound since very few of us had access pre-launch.

The video I posted referenced this.

In summary: The person had access to early releases through his work at Microsoft Research where they were integrating GPT-4 into Bing. He used "Draw a unicorn in TikZ" (TikZ is probably the most complex and powerful tool to create graphic elements in LaTeX) as a prompt and noticed how the model's responses changed with each release they got from OpenAI. While at first the drawings got better and better, once OpenAI started focusing on "safety" subsequent releases got worse and worse at the task.

Post reply on HN