Live data from Hacker News

Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

news.ycombinator.com

281–290 of 817 posts

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#282

There's no doubt that it's gotten a lot worse on coding, I've been using this benchmark on each new version of GPT-4 "Write a tiptap extension that toggles classes" and so far it's gotten it right every time, but not any more, now it hallucinates a simplified solution that don't even use the tiptap api any more. It's also 200% more verbose in explaining it's reasoning, even if that reasoning makes no sense whatsoever…

Wait so you’ve gotten GPT 4 to successfully write TipTap extensions for you? Are you using Copilot or the ChatGPT app?

Not only writing, extending and figuring out quite complicated usage based on the API documentation. I'll open source some of them in the near future. I'm using ChatGPT Plus with GPT-4, that gave the best results. Also worked via API key and custom prompts.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#283

OpenAI's models feel 100% nerfed to me at this point. I had it solving incredibly complex problems a few months ago (i.e. write a minimal PDF parser example), but today you will get scolded for asking such a complicated task of it. I think they programmed a classifier layer to detect certain coding tasks and shut it down with canned BS. I like to imagine certain billion/trillion-dollar mega corps had a back-room say…

So far my experience with Vicunlocked30b has been pleasant. https://huggingface.co/TheBloke/VicUnlocked-30B-LoRA-GGML Although I haven't had much of my time available for this recently. My recommendation would be to start with https://github.com/oobabooga/text-generation-webui You will find almost everything you need to know there and on 4chan.org/g/catalog - search for LMG.

You should beware that /lmg/ is full of horrible people, discussing horrible things, like most of 4chan. Reddit's r/locallama is much more agreeable. That said, the 4chan thread tends to be more up-to-date. These guys are serious about their ERP.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#284
post #275

Earlier quoted context omitted.

It hasn't been changed since March 14th... So it's equally nerfed as it was then... Also, the playground lets you set the 'system message', which you can use to tell it to answer questions even if the results may be dangerous/rude/inappropriate.

I have been using the API. There are conflicting reports in this thread that seems to indicate it may also be affected. I am not sure.

Did you set the model code to gpt-4-0314?

I did, and I still get the original speed (produces tokens at about the speed you would read aloud), and I haven't seen quality change.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#285

My guess is that -probably no. It's more likely you had a stream of good luck in your earlier interactions and now you're observing regression to the mean. That can easily happen and it's why, for example, medical studies, are not taken as definitive proof of an effect. To further clarify, regression to the mean is the inevitable consequence of statistical error. Suppose (classic example) we want to test a hypertensi…

How do you explain people issuing the same prompt over time as a test and getting worse and worse responses?

Well because it's just what the parent said, it's all a subjective experience, and maybe the anthropomorphism element to it blew people away more than the actually content of the responses ? Ie, you're just used to it now.

The human mind is ridiculously fickle, it takes a lot to be impressed for more than a few days / weeks.

It did seem radically cool at first but over time I got quite sick of using it too.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#286
post #257

Earlier quoted context omitted.

They're up against a pretty difficult barrier - if we had a perfect all-knowing oracle it might easily have opinions that are racist. Statistics alone suggest there will be racist truths. We're dealing with groups of people who are observably different from each other in correlated ways. GPT would need to reach a convincing balance of lying and honesty if it is supposed to navigate that challenge. It'd have to be dee…

Can you expand on the last sentence of your first paragraph?

Crime stats, average IQ across groups, stereotype accuracy, etc.

What's interesting to me is not the above, which is naughty in the anglosphere, but the question of the unknown unknowns that could be as bad or worse in other cultural contexts. There are probably enough people of Indian descent involved in GPT's development that they could guide it past some of the caste landmines, but what about a country like Turkey? We know they have massive internal divisions, but do we know what would exacerbate them and how to avoid them? What about Iran, or South Africa, or Brazil?

We RLHF the piss out of LLMs to ensure they don't say things that make white college graduates in San Francisco ornery, but I'd suggest the much greater risk lies in accidentally spawning scissor statements in cultures you don't know how to begin to parse to figure out what to avoid.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#287
post #207
post #200

Earlier quoted context omitted.

Same here. If I have a choice between honesty and political correctness, I always pick honesty.

It's not about honesty vs. political correctness, it is about safety. There's real concern that the model can cause harm to humans, in a variety of ways, which is and should be unethical. If we have to argue about that in 2023, that's concerning.

[dead]

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#288

OpenAI's models feel 100% nerfed to me at this point. I had it solving incredibly complex problems a few months ago (i.e. write a minimal PDF parser example), but today you will get scolded for asking such a complicated task of it. I think they programmed a classifier layer to detect certain coding tasks and shut it down with canned BS. I like to imagine certain billion/trillion-dollar mega corps had a back-room say…

I agree if this trend continues even inferior local models are going to have value just because the public apis are so limited.

> Conspiracy shenanigans aside, I've decided to cancel my "premium" membership and am exploring open/DIY models.

The crazy thing is that this is an application that really benefits from being in the cloud because the high vram gpus are so expensive that it makes sense to batch requests from many users to maximize utilization.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#289
Reading the comments in this thread, with the rightful distrust of OpenAI and criticism of the model, it occurs to me that the underlying problem we’re facing here comes down to stakeholders and incentive structures.

Ai will not be a technical problem (nor a solution!), rather our civilization will continue to be bottlenecked by problems of culture. OpenAI will succeed/fail for cultural reasons, not technical ones. Humanity will benefit from or be harmed by ai for cultural reasons, not technical ones.

I don’t have any answers here. I do however get the continuous impression that ai has not shifted the ground under our feet. The bottlenecks and underlying paradigms remain basically the same.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#290

Given the incoming compute capability from nvidia and the speed of advancement, we have to stop and think ... does it make sense to give access, paid or otherwise, to these models once they reach a certain sophistication? Or does it make even more sense to hoard the capability to out compete any competitor, of any kind, commercially or politically and hide the true extent of your capability to avoid scrutiny and legi…

Perhaps this could explain Sam Altman's unique arrangement with OpenAI?

I.e. "Give me unfettered access to the latest models (especially the secret ones) and you can keep your money"

Post reply on HN