Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
11–20 of 817 posts
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#12One of my least favorite patterns that tech companies do is use “Experimentation” overzealously or prematurely. Mainly, my problem is they’re not transparent about it, and it creates an inconsistent product experience that just confuses you - why did this one Zillow listing have this UI order but the similar one I clicked seconds later had a different one? Why did this page load on Reddit get some weirdass font? Because it’s an experiment the bar to launch is low and you’re not gonna find any official blog posts about the changes until it’s official. And when it causes serious problems, there’s nowhere to submit a form or tell you why, and only very rarely would support, others, or documentation even realize some change is from an experiment. Over the past few years I’ve started noticing this everywhere online.
Non-sticky UI experiments are especially bad because at eg 1% of pageloads the signal is going to be measuring users asking themselves wtf is up and temporarily spending more time on page trying to figure out where the data moved. Sticky and/or less noticeable experiments like what this could be have stronger signals but are even more annoying as a user, because there’s no notice that you’re essentially running some jank beta version, and no way to opt back into the default - for you it’s just broken. Especially not cool if you’re a paying customer.
I’m not saying it’s necessarily an experiment, it could be just a regular release or nothing at all. I’d hope if OpenAI was actually reducing the parameter size of their models they’d publicly announce that, but I could totally see them running an experiment measuring how a cheaper, smaller model affects usage and retention without publishing anything, because it’s exactly the kind of “right hand doesn’t know what the left is doing” thing that happens at fancy schmancy tech companies.
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#13Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#14The original GPT-4 felt like magic to me, I had this sense of awe while interacting with it. Now it is just a dumb stochastic parrot.
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#15Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#16Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#17Is it consistently worse or just sometimes/often worse than before? Any extreme power users or GPT-whisperers here? If it’s only noticeably worse X% of the time my bet would be experimentation. One of my least favorite patterns that tech companies do is use “Experimentation” overzealously or prematurely. Mainly, my problem is they’re not transparent about it, and it creates an inconsistent product experience that jus…
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#18Before, if I had an issue with a library or debugging issue, it would try to be helpful and walk me through potential issues, and ask me to 'let it know' if it worked or not. Now it will try to superficially diagnose the problem and then ask me to check the online community for help or continuously refer me to the maintainers rather than trying to figure it out.
Similarly, I had been using it to help me think through problems and issues from different perspectives (both business and personal) and it would take me in-depth through these. Now, again, it gives superficial answers and encourages going to external sources.
I think if you keep pressing in the right ways it'll eventually give in and help you as it did before, but I guess this will take quite a bit of prompting.
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#19I feel the same way. It feels…lazy now.
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#20Yes! It didn't even try on my question of Jarvis standings desks, which is a fairly old product that hasn't changed up.. Their typical "My knowledge cutoff..." response doesn't even make sense. It screwed up another question I asked it about server uptime and four-9s, Bard got it right. I've moved back to Bard for the time being...It's way faster as well. And GPT-4's knowledge cutoff thing is getting old fast. Exampl…