Live data from Hacker News

Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

news.ycombinator.com

271–280 of 817 posts

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#271

Do we have a good, objective benchmark set of prompts in existence somewhere? If not, I think having one would really help with tracking changes like that. I'm always skeptical of subjective feelings of tough-to-quantify things getting worse or better, especially where there is as much hype as for the various AI models. One explanation for the feelings is the model really getting significantly worse over time. Anothe…

https://github.com/FranxYao/chain-of-thought-hub

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#272

Earlier quoted context omitted.

It’s go-to tactic now if I ask it to go over any piece of code is to give a generic overview. Earlier, it would section out the code into chunks and go through each one individually.

Yeah, the bing integration did not go well. Went from amazing to annoying.

Aren’t the original weights around somewhere?

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#273

There's no doubt that it's gotten a lot worse on coding, I've been using this benchmark on each new version of GPT-4 "Write a tiptap extension that toggles classes" and so far it's gotten it right every time, but not any more, now it hallucinates a simplified solution that don't even use the tiptap api any more. It's also 200% more verbose in explaining it's reasoning, even if that reasoning makes no sense whatsoever…

Wait so you’ve gotten GPT 4 to successfully write TipTap extensions for you? Are you using Copilot or the ChatGPT app?

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#274

My guess is that -probably no. It's more likely you had a stream of good luck in your earlier interactions and now you're observing regression to the mean. That can easily happen and it's why, for example, medical studies, are not taken as definitive proof of an effect. To further clarify, regression to the mean is the inevitable consequence of statistical error. Suppose (classic example) we want to test a hypertensi…

How do you explain people issuing the same prompt over time as a test and getting worse and worse responses?

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#275
post #251

Earlier quoted context omitted.

So are you sure that isn't also nerfed?

It hasn't been changed since March 14th... So it's equally nerfed as it was then... Also, the playground lets you set the 'system message', which you can use to tell it to answer questions even if the results may be dangerous/rude/inappropriate.

I have been using the API. There are conflicting reports in this thread that seems to indicate it may also be affected. I am not sure.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#277

OpenAI's models feel 100% nerfed to me at this point. I had it solving incredibly complex problems a few months ago (i.e. write a minimal PDF parser example), but today you will get scolded for asking such a complicated task of it. I think they programmed a classifier layer to detect certain coding tasks and shut it down with canned BS. I like to imagine certain billion/trillion-dollar mega corps had a back-room say…

So far my experience with Vicunlocked30b has been pleasant. https://huggingface.co/TheBloke/VicUnlocked-30B-LoRA-GGML

Although I haven't had much of my time available for this recently. My recommendation would be to start with https://github.com/oobabooga/text-generation-webui

You will find almost everything you need to know there and on 4chan.org/g/catalog - search for LMG.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#278
post #14

Yes. Before the update, when its avatar was still black, it solved pretty complex coding problems effortlessly and gave very nuanced, thoughtful answers to non-programming questions. Now it struggles with just changing two lines in a 10-line block of CSS and printing this modified 10-line block again. Some lines are missing, others are completely different for no reason. I'm sure scaling the model is hard, but they l…

"The original GPT-4 felt like magic to me" You never had access to that original. Watch this talk by one of the people that integrated GPT-4 in Bing telling how they noticed GPT-4 releases they got from OpenAI got iteratively and significantly nerfed even during the project. https://www.youtube.com/watch?v=qbIk7-JPB2c

Here's another interview from a guy who had access to the unfiltered GPT-4 before its release. He says it was extremely powerful and would answer any question whatsoever without hesitating.

https://www.youtube.com/watch?v=oLiheMQayNE&t=2849s

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#279
post #257
post #124

The reason it's worse is basically because it's more 'safe' (not racist, etc). That of course sounds insane, and doesn't mean that safety shouldn't be strived for, etc - but there's an explanation as to how this occurs. It occurs because the system essentially does a latent classification of problems into 'acceptable' or 'not acceptable' to respond to. When this is done, a decent amount of information is lost regardi…

They're up against a pretty difficult barrier - if we had a perfect all-knowing oracle it might easily have opinions that are racist. Statistics alone suggest there will be racist truths. We're dealing with groups of people who are observably different from each other in correlated ways. GPT would need to reach a convincing balance of lying and honesty if it is supposed to navigate that challenge. It'd have to be dee…

How is racism different from stereotype?

How is stereotype different from pattern recognition?

These questions don't seem to go through the minds of people when developing "unbiased/impartial" technology.

There is no such thing as objective. So, why pretend to be objective and unbiased, when we all know its a lie?

Worst, if you pretend to be objective but aren't, then you are actually racist.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#280

Earlier quoted context omitted.

Try out Bard, it's coding is much improved in the last 2 weeks. I've unfortunately switched over for the time being.

No thanks! I have better things to do than feeding that advertising behemoth. What I like about ChatGPT is that I don't see any ads at all!

That you know of.

Don't you worry, if there is any medium, place or mode of interaction people spend time on, advertising will eventually metastasize to it, and will keep growing until it completely devalues the activity and destroys most of the utility it provides.

Post reply on HN