Live data from Hacker News

Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

news.ycombinator.com

71–80 of 817 posts

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#71
post #14

Yes. Before the update, when its avatar was still black, it solved pretty complex coding problems effortlessly and gave very nuanced, thoughtful answers to non-programming questions. Now it struggles with just changing two lines in a 10-line block of CSS and printing this modified 10-line block again. Some lines are missing, others are completely different for no reason. I'm sure scaling the model is hard, but they l…

"The original GPT-4 felt like magic to me" You never had access to that original. Watch this talk by one of the people that integrated GPT-4 in Bing telling how they noticed GPT-4 releases they got from OpenAI got iteratively and significantly nerfed even during the project. https://www.youtube.com/watch?v=qbIk7-JPB2c

“You never had access to that original.”

While your overall point is well taken, GP is clearly referring to the original public release of GPT-4 on March 14.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#72
post #4

Yes! It didn't even try on my question of Jarvis standings desks, which is a fairly old product that hasn't changed up.. Their typical "My knowledge cutoff..." response doesn't even make sense. It screwed up another question I asked it about server uptime and four-9s, Bard got it right. I've moved back to Bard for the time being...It's way faster as well. And GPT-4's knowledge cutoff thing is getting old fast. Exampl…

According to https://ukstore.hermanmiller.com/products/jarvis-bamboo-desk... Bard is off by about 10%, but generally correct (it's also possible that the US dimensions are actually smaller and could account for the difference).

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#73

Earlier quoted context omitted.

Same for me, I’m in Estonia :(

You can use a VPN to use an American connection, it doesn't matter where your Google account is registered.

Not necessarily American, you just have to avoid EU and, I believe, Russia/China/Cuba etc.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#74

There's no doubt that it's gotten a lot worse on coding, I've been using this benchmark on each new version of GPT-4 "Write a tiptap extension that toggles classes" and so far it's gotten it right every time, but not any more, now it hallucinates a simplified solution that don't even use the tiptap api any more. It's also 200% more verbose in explaining it's reasoning, even if that reasoning makes no sense whatsoever…

Do you have API access? If so, have you tried your tiptap question on the gpt-4-0314 model? That is supposedly the original version released to the public on March 14.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#75
I had the same feeling with GTP3.5 yesterday, I've asked if in order to calculate ARPA you need to consider the free tier and it come out with some gibberish about the fact that he doesn't know anything post 2021 about the Advanced Research Projects Agency.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#76
It's time to design a public benchmark for these types of systems to compare between versions. Of course, any vendor who trains on the benchmark should face extreme contempt, but we'd also need to generate novel questions of equal complexity.

Alternatively, there should be a trusted auditor who uses a secret benchmark.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#77

To me, it feels like it's started giving superficial responses and encouraging follow-up elsewhere -- I wouldn't be surprized if its prompt has changed to something to that effect. Before, if I had an issue with a library or debugging issue, it would try to be helpful and walk me through potential issues, and ask me to 'let it know' if it worked or not. Now it will try to superficially diagnose the problem and then a…

[dead]

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#78

There's no doubt that it's gotten a lot worse on coding, I've been using this benchmark on each new version of GPT-4 "Write a tiptap extension that toggles classes" and so far it's gotten it right every time, but not any more, now it hallucinates a simplified solution that don't even use the tiptap api any more. It's also 200% more verbose in explaining it's reasoning, even if that reasoning makes no sense whatsoever…

It was a great ride while it lasted. My assumption is that efficacy at coding tasks is such a small percent of users, they’ve just sacrificed it on the altar of efficiency and/or scale. That, or they’ve cut some back room deal with Microsoft to make Copilot have access to the only version of the model that can actually code.

Copilot X (the new version, with a chat interface etc) is significantly worse than GPT-4 (at least before this update). It felt like gpt3.5-turbo to me.

Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?

#80

There's no doubt that it's gotten a lot worse on coding, I've been using this benchmark on each new version of GPT-4 "Write a tiptap extension that toggles classes" and so far it's gotten it right every time, but not any more, now it hallucinates a simplified solution that don't even use the tiptap api any more. It's also 200% more verbose in explaining it's reasoning, even if that reasoning makes no sense whatsoever…

It was a great ride while it lasted. My assumption is that efficacy at coding tasks is such a small percent of users, they’ve just sacrificed it on the altar of efficiency and/or scale. That, or they’ve cut some back room deal with Microsoft to make Copilot have access to the only version of the model that can actually code.

Honestly, why not different versions at this point? People who want it for coding don't care if it knows the history of prerevolution France, and vice versa.

Seems they could wow more people if they had specialized versions, rather than the jack of all trades that tries to exist now.

Edit: Oh God, I just described our human system of specialty and how the AI could replace us using the same means...

Post reply on HN